Fine adjustment method and device for large model in computer hardware operation and maintenance, equipment and medium
By decomposing the weight matrix of the pre-trained large model into the main space and residual space, and performing targeted fine-tuning, the problem of weak performance of traditional large models in decision-making functional scenarios is solved, and efficient functional enhancement and training efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510724514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-26
AI Technical Summary
Traditional large models perform weakly in decision-making functions and other scenarios where causal logical reasoning, rule constraints or multi-objective optimization, and the full parameter fine-tuning method leads to large training overhead and is prone to overfitting or pre-training knowledge forgetting.
By decomposing the weight matrix of the pre-trained large model into the main space and the residual space, the weight matrix corresponding to the original function and the target incremental function is adjusted respectively, and the matrix decomposition and low-rank matrix initialization are used to perform targeted fine-tuning to maintain the original function and enhance the decision-making function.
It realizes that while reducing training overhead and avoiding overfitting, the decision-making function capabilities of the large model are improved, ensuring the stability of the original function and the learning efficiency of the target incremental function.
Smart Images

Figure CN120542601A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and computer technology, and in particular to a method, apparatus, device, and medium for fine-tuning a large model in computer hardware operation and maintenance. Background Art
[0002] Traditional large models typically focus on tasks that rely on the statistical laws of structured data, such as recommendation functions and pattern recognition, but are relatively weak in scenarios such as decision-making that require causal logical reasoning, rule constraints, or multi-objective optimization. With the continuous evolution of large model capabilities, it has become possible to optimize their decision-making capabilities through fine-tuning techniques. In the field of intelligent computing operations, using large model fine-tuning technology to adapt to downstream tasks has become a trend. However, related technologies generally use full parameter fine-tuning, which has high training overhead, easy to cause overfitting, or forget pre-training knowledge, resulting in poor training results. Summary of the Invention
[0003] One aspect of the present disclosure provides a method for fine-tuning a large model in computer hardware operation and maintenance, comprising: determining an adjustable weight matrix that needs to be fine-tuned for a pre-trained large model, the adjustable weight matrix comprising a first weight matrix corresponding to a main space of a feature space of the pre-trained large model, and a second weight matrix corresponding to a residual space of the feature space of the pre-trained large model, the main space representing the feature space where the functional features of the original functions of the pre-trained large model are located, the residual space representing the feature space where the functional features of the target incremental functions are located, the original functions comprising a recommendation function and a pattern recognition function corresponding to computer hardware, the target incremental functions comprising a decision function corresponding to computer hardware, the matrix parameters in the main space and the residual space corresponding to computer hardware parameters ... target incremental functions comprising a decision function corresponding to computer hardware, the target incremental functions comprising a recommendation function and a pattern recognition function corresponding to computer hardware, the target incremental functions comprising a decision function corresponding to computer hardware, the target incremental functions comprising a recommendation function and a decision function The parameters in the feature space correspond to the recommended function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function correspond to the decision function parameters of the computer hardware; and according to the target incremental function, the weight values in the first weight matrix and the second weight matrix are adjusted to obtain the adjusted first weight matrix and the adjusted second weight matrix, wherein the adjustment of the second weight matrix is based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space, the adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
[0004] According to an embodiment of the present disclosure, the fine-tuning method of a large model in computer hardware operation and maintenance also includes: performing matrix decomposition on an adjustable weight matrix, determining a source basis of the adjustable weight matrix in the source space and a target basis in the target space, wherein the source basis is used to map the data input into the pre-trained large model from the source feature space to the target feature space, and the target basis is used to map the data that has completed data transformation in the target space back to the source feature space, the source basis includes multiple source vectors, and the target basis includes multiple target vectors; according to the number of feature information included in the source vector, a key source basis is formed by a preset number of source vectors in the source basis, and a supplementary source basis is formed by other source vectors; according to the number of feature information included in the target vector, a key target basis is formed by a preset number of target vectors in the target basis, and a supplementary target basis is formed by other target vectors; based on the key source basis and the key target basis, the main space is determined, and based on the supplementary source basis and the supplementary target basis, the residual space is determined.
[0005] According to an embodiment of the present disclosure, the fine-tuning method for large models in computer hardware operation and maintenance also includes: determining a first weight matrix based on a key source basis, a key target basis and a first training matrix, wherein the first training matrix is determined based on a first low-rank matrix and a second low-rank matrix, the first low-rank matrix and the second low-rank matrix are obtained by random initialization, the number of rows of the first low-rank matrix is determined according to the number of columns of the key source basis, the number of columns of the second low-rank matrix is determined according to the number of rows of the key target basis, and the number of columns of the first low-rank matrix is the same as the number of rows of the second low-rank matrix; and determining a second weight matrix based on a supplementary source basis, a supplementary target basis and the second training matrix, wherein the second training matrix is determined based on a third low-rank matrix and a fourth low-rank matrix, the third low-rank matrix and the fourth low-rank matrix are obtained by random initialization, the number of rows of the third low-rank matrix is determined according to the number of columns of the supplementary source basis, the number of columns of the fourth low-rank matrix is determined according to the number of rows of the supplementary target basis, and the number of columns of the third low-rank matrix is the same as the number of rows of the fourth low-rank matrix.
[0006] According to an embodiment of the present disclosure, a first weight matrix is determined based on a key source basis, a key target basis and a first training matrix, including: determining the feature direction in the main space based on the key source basis and the first low-rank matrix; determining the feature quantity in the main space based on the key target basis and the second low-rank matrix; and determining the first weight matrix based on the feature direction in the main space and the feature quantity in the main space.
[0007] According to an embodiment of the present disclosure, the weight matrix of the pre-trained large model includes an adjustable weight matrix and a fixed weight matrix; according to the target incremental function, the weight values in the first weight matrix and the second weight matrix are adjusted to obtain the adjusted first weight matrix and the adjusted second weight matrix, including: initializing the second low-rank matrix to a preset initial value so that the feature quantity in the initialized main space obtained after initialization meets the preset conditions; determining the initialization result of the first weight matrix based on the feature quantity in the initialized main space and the feature direction in the main space; processing the data input to the pre-trained large model based on the fixed weight matrix and the initialized first weight matrix to obtain output data; determining the model loss based on the output data and the expected output data; and adjusting the parameters of the first low-rank matrix and the second low-rank matrix to achieve parameter adjustment of the first weight matrix until the model loss converges.
[0008] According to an embodiment of the present disclosure, the fine-tuning method of a large model in computer hardware operation and maintenance also includes: determining a first learning rate of a first weight matrix based on a preset learning rate reference value and an eigenvalue of a main space; determining a second learning rate of a second weight matrix based on a preset learning rate reference value and an eigenvalue of a residual space; and training the first training matrix using the first learning rate and training the second training matrix using the second learning rate.
[0009] According to an embodiment of the present disclosure, the fine-tuning method for large models in computer hardware operation and maintenance also includes: determining a first parameter amount to be adjusted in the main space based on a first training matrix; determining a second parameter amount to be adjusted in the residual space based on a second training matrix; adjusting the first norm ratio based on the first parameter amount to be adjusted and the second parameter amount to be adjusted to obtain an adjusted first norm ratio; and regularizing the main space and the residual space respectively based on the first parameter amount to be adjusted, the second parameter amount to be adjusted, the adjusted first norm ratio and the second norm ratio.
[0010] Another aspect of the present disclosure provides a method and apparatus for fine-tuning a large model in computer hardware operation and maintenance, comprising: a matrix determination module for determining an adjustable weight matrix that needs to be fine-tuned for a pre-trained large model, the adjustable weight matrix comprising a first weight matrix corresponding to a main space of a feature space of the pre-trained large model, and a second weight matrix corresponding to a residual space of the feature space of the pre-trained large model, the main space representing the feature space where the functional features of the original functions of the pre-trained large model are located, the residual space representing the feature space where the functional features of the target incremental functions are located, the original functions comprising a recommendation function and a pattern recognition function corresponding to computer hardware, the target incremental functions comprising a decision function corresponding to computer hardware, the matrix parameters in the main space and the residual space corresponding to computer hardware parameters, the original functions The parameters in the feature space correspond to the recommended function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function correspond to the decision function parameters of the computer hardware; and a matrix adjustment module is used to adjust the weight values in the first weight matrix and the second weight matrix according to the target incremental function to obtain the adjusted first weight matrix and the adjusted second weight matrix, wherein the adjustment of the second weight matrix is based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space, the adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
[0011] Another aspect of the present disclosure provides a non-volatile storage medium storing computer-executable instructions, wherein the instructions are used to implement any of the above methods when executed.
[0012] Another aspect of the present disclosure provides a computer program, comprising computer-executable instructions, wherein the instructions are used to implement any of the above methods when executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] For a more complete understanding of the present disclosure and its advantages, reference will now be made to the following description taken in conjunction with the accompanying drawings, in which:
[0014] Figure 1 The application scenarios of the method, apparatus, device, and medium for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure are schematically illustrated;
[0015] Figure 2 A flowchart of a method for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure is schematically shown;
[0016] Figure 3ASchematically shows a parameter distribution diagram in the main space obtained by dividing the fine-tuning method of a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure;
[0017] Figure 3B Schematically shows a parameter distribution diagram in the residual space obtained by dividing the fine-tuning method of a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure;
[0018] Figure 4 The figure schematically shows a flow chart of initializing the weight matrix of a pre-trained large model according to a method for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure;
[0019] Figure 5 A block diagram schematically illustrates a device for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure; and
[0020] Figure 6 A schematic block diagram of an example electronic device that can be used to implement the method of the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely illustrative and are not intended to limit the scope of the present disclosure. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0024] The accompanying drawings show some block diagrams and / or flow charts. It should be understood that some blocks in the block diagrams and / or flow charts, or combinations thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when these instructions are executed by the processor, they can create a device for implementing the functions / operations described in the block diagrams and / or flow charts.
[0025] Therefore, the techniques of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). In addition, the techniques of the present disclosure can take the form of a computer program product on a computer-readable medium having stored thereon instructions, which can be used by or in conjunction with an instruction execution system. In the context of the present disclosure, a computer-readable medium can be any medium that can contain, store, convey, propagate, or transmit instructions. For example, a computer-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. Specific examples of computer-readable media include: magnetic storage devices, such as magnetic tape or hard disk drives (HDDs); optical storage devices, such as compact disks (CD-ROMs); memory, such as random access memory (RAM) or flash memory; and / or wired or wireless communication links.
[0026] An embodiment of the present disclosure provides a method for fine-tuning a large model in computer hardware operation and maintenance, comprising: determining an adjustable weight matrix that needs to be fine-tuned for a pre-trained large model, the adjustable weight matrix comprising a first weight matrix corresponding to the main space of the feature space of the pre-trained large model, and a second weight matrix corresponding to the residual space of the feature space of the pre-trained large model, the main space representing the feature space where the functional features of the original function of the pre-trained large model are located, the residual space representing the feature space where the functional features of the target incremental function are located, the matrix parameters in the main space and the residual space corresponding to the computer hardware parameters, the parameters in the feature space of the original function corresponding to the recommended functions and pattern recognition of the computer hardware The parameters in the feature space of the target incremental function correspond to the decision function parameters of the computer hardware; and according to the target incremental function, the weight values in the first weight matrix and the second weight matrix are adjusted to obtain the adjusted first weight matrix and the adjusted second weight matrix, wherein the adjustment of the second weight matrix is based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space, the adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
[0027] Figure 1 The application scenarios of the method, apparatus, device and medium for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure are schematically illustrated.
[0028] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0030] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0031] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0032] It should be noted that the method for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the apparatus for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0033] Alternatively, the method for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiment of the present disclosure may also generally be executed by the server 105. Accordingly, the device for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiment of the present disclosure may generally be provided in the server 105. The method for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiment of the present disclosure may also be executed by a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the device for fine-tuning a large model in computer hardware operation and maintenance provided by the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] Figure 2 The flowchart of the method for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure is schematically shown.
[0036] Specifically, if Figure 2 As shown, the method includes operations S210 to S220.
[0037] In operation S210 , an adjustable weight matrix that needs to be fine-tuned for the pre-trained large model is determined.
[0038] According to the embodiments of the present disclosure, the pre-trained big model is a big model that has mastered a knowledge base composed of part of the required knowledge through pre-training and has certain capabilities or functions. The capabilities or functions that the pre-trained big model has can be called the original functions of the pre-trained big model.
[0039] According to the embodiments of the present disclosure, due to business expansion and other reasons, it is necessary to enable the big model to have more and more comprehensive functions. The functions that need to be newly mastered by the big model can be called the target incremental functions of the pre-trained big model.
[0040] Typically, large language models, thanks to their deep neural network architecture, are able to automatically learn and extract rich, multi-layered features from massive amounts of data. Therefore, most large language models are quite effective at tasks such as recommendation and pattern recognition. However, decision-making tasks are inherently complex and deeply nonlinear. Therefore, in scenarios requiring clear decision logic and high interpretability, the decision results of large language models are often difficult to trust and apply. Therefore, the original functions of pre-trained large models include recommendation and pattern recognition functions corresponding to computer hardware, and the target incremental functions of pre-trained large models include decision-making functions corresponding to computer hardware.
[0041] According to the embodiments of the present disclosure, different target incremental functions usually require the pre-trained large model to have different knowledge bases. If the knowledge base that the pre-trained large model has mastered is different from the knowledge base required for the target incremental function, the pre-trained large model is also required to expand the knowledge base during the training process of the pre-trained large model, which results in extremely high training costs. Therefore, according to the target incremental function, a pre-trained large model that has mastered the knowledge base required for the target incremental function can be selected, and the pre-trained large model can be subsequently trained to ensure that the training process is as simple as possible and to control the training cost.
[0042] According to an embodiment of the present disclosure, the adjustable weight matrix includes a first weight matrix corresponding to the main space of the feature space of the pre-trained large model, and a second weight matrix corresponding to the residual space of the feature space of the pre-trained large model, wherein the main space represents the feature space in which the functional characteristics of the original function of the pre-trained large model are located, and the residual space represents the feature space in which the functional characteristics of the target incremental function are located. The functional characteristics of the original function and the functional characteristics of the target incremental function are different.
[0043] According to an embodiment of the present disclosure, the matrix parameters in the main space and the residual space correspond to computer hardware parameters, the parameters in the feature space of the original function correspond to the recommendation function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function correspond to the decision function parameters of the computer hardware.
[0044] In operation S220 , weight values in the first weight matrix and the second weight matrix are adjusted according to the target increment function to obtain an adjusted first weight matrix and an adjusted second weight matrix.
[0045] According to an embodiment of the present disclosure, the adjustment of the second weight matrix is performed based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is performed based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space. The adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
[0046] According to an embodiment of the present disclosure, the second weight matrix can be adjusted according to the data set of the target incremental function, and the first weight matrix can be adjusted according to the data set of the original function. Since the purpose of setting the first weight matrix is to ensure that the adjusted large model can retain the original function, the data set of the original function can be directly output using the pre-trained large model.
[0047] According to an embodiment of the present disclosure, since the adjustable weight matrix will also change during the process of adjusting the weight values in the second weight matrix according to the target incremental function, this will cause the pre-trained large model to perform the original function less effectively after the parameter adjustment is completed. Therefore, after adjusting the weight values in the second weight matrix according to the target incremental function, the weight values in the first weight matrix can be adjusted using the original function, thereby ensuring that the pre-trained large model still has the ability to perform the original function after the parameter adjustment is completed.
[0048] According to an embodiment of the present disclosure, when performing incremental learning on a pre-trained large model, the adjustable weight matrix that needs to be fine-tuned is decomposed into a first weight matrix and a second weight matrix, and the first weight matrix corresponds to the main space in the feature space of the pre-trained large model, and the second weight matrix corresponds to the residual space in the feature space of the pre-trained large model. During the training process, the first weight matrix is used to maintain the original function of the pre-trained large model, and the second weight matrix is used to learn the target incremental function of the pre-trained large model. Since the original function and the target incremental function require the same model parameters and knowledge base reserves, after using the first weight matrix to train and maintain the original function, when adjusting the weight values in the first weight matrix and the second weight matrix according to the target incremental function, there is no need to learn the knowledge base reserves, and incremental learning can be completed more quickly and conveniently.
[0049] According to an embodiment of the present disclosure, the fine-tuning method of a large model in computer hardware operation and maintenance also includes: performing matrix decomposition on an adjustable weight matrix, determining a source basis of the adjustable weight matrix in the source space and a target basis in the target space, wherein the source basis is used to map the data input into the pre-trained large model from the source feature space to the target feature space, and the target basis is used to map the data that has completed data transformation in the target space back to the source feature space, the source basis includes multiple source vectors, and the target basis includes multiple target vectors; according to the number of feature information included in the source vector, a key source basis is formed by a preset number of source vectors in the source basis, and a supplementary source basis is formed by other source vectors; according to the number of feature information included in the target vector, a key target basis is formed by a preset number of target vectors in the target basis, and a supplementary target basis is formed by other target vectors; based on the key source basis and the key target basis, the main space is determined, and based on the supplementary source basis and the supplementary target basis, the residual space is determined.
[0050] According to an embodiment of the present disclosure, the source space and the target space are obtained by performing matrix decomposition on the adjustable weight matrix, as shown in formula (1):
[0051] (1)
[0052] Among them, W is the m×n dimensional adjustable weight matrix, U is the matrix corresponding to the m×m dimensional source space, V T is the matrix corresponding to the target space of n×n dimensions, S is the m×n dimensional diagonal matrix, ,in, are the singular values of the matrix W arranged in descending order, d is the rank of the matrix S, d=min{m,n}.
[0053] According to an embodiment of the present disclosure, the source basis is a basis constituting a source space, the source basis includes multiple source basis vectors, and the number of components of each source basis vector and the number of source basis vectors included in the source basis are the same as the spatial dimension of the source space.
[0054] According to an embodiment of the present disclosure, the target basis is a basis that constitutes the target space, and the target basis includes multiple target basis vectors. The number of components of each target basis vector and the number of target basis vectors included in the target basis are the same as the spatial dimension of the target space.
[0055] According to an embodiment of the present disclosure, the matrix U corresponding to the source space is composed of d source vectors, and the matrix V corresponding to the target space is T It is composed of d target vectors, and both the source vector and the target vector are singular vectors, where U=[ ], V T Transpose to get V=[ ].
[0056] According to an embodiment of the present disclosure, based on a preset number k, the k source vectors including the largest number of feature information among the d source vectors can be determined as the key source basis U according to the number of feature information included in each of the d source vectors. p =[ ], and determine the other dk source vectors as the supplementary source basis U res =[ ].
[0057] Similarly, according to the number of characteristic information included in each of the d target vectors, the k target vectors with the largest number of characteristic information among the d target vectors can be determined as the key target basis V p =[ ], and determine the other dk target vectors as the supplementary target basis V res =[ ].
[0058] According to embodiments of the present disclosure, since the key source basis and the key target basis are respectively composed of singular vectors in the source basis and the target basis that include more characteristic information, the main space spanned by the key source basis and the key target basis can include more characteristic information in the adjustable weight matrix. Similarly, the residual space spanned by the supplementary source basis and the supplementary target basis can include less characteristic information in the adjustable weight matrix.
[0059] According to the embodiments of the present disclosure, since the pre-trained large model has already learned the knowledge base required to execute the original function, and the knowledge base required by the pre-trained large model to execute the target incremental function is the same as the knowledge base required to execute the original function, it can be determined that during the process of incremental function training, the pre-trained large model only needs to train the capabilities required to drive the target incremental function. Furthermore, since incremental function training actually fine-tunes the adjustable weight matrix, it can be determined that the residual space is used to learn the target incremental function, and the main space is used to maintain the original function.
[0060] According to an embodiment of the present disclosure, by further dividing the matrix decomposition result of the adjustable weight matrix, a key source basis, a supplementary source basis, a key target basis and a supplementary target basis are obtained, and the main space and the residual space are further determined. Since the main space is determined based on the key source basis and the key target basis, and the key source basis and the key target basis are respectively composed of vectors with a large number of feature information included in the source basis and the target basis, the main space relatively includes more features of the adjustable weight matrix, and more features can be used to maintain the original function of the pre-trained large model, thereby ensuring that the pre-trained large model can still retain the original function after incremental learning. Relatively speaking, the residual space relatively includes a small number of features in the adjustable weight matrix, and using a small number of features to perform incremental learning on the pre-trained large model can reduce the impact on the original function of the pre-trained large model while ensuring that the target incremental function is learned.
[0061] Figure 3A The diagram schematically shows the parameter distribution in the main space obtained by dividing the large model fine-tuning method in computer hardware operation and maintenance according to an embodiment of the present disclosure.
[0062] like Figure 3A As shown in the figure, the horizontal and vertical coordinates represent the new feature dimensions obtained after the bit data is reduced in dimension. The markers of different shapes in the coordinates correspond to different features in the main space. Figure 3A It can be seen that except for a small number of markers scattered around the periphery, most of the other markers are at the (0, 0) coordinate. The feature distribution is relatively concentrated, which is more suitable for maintaining the original function of the pre-trained large model.
[0063] Figure 3B The diagram schematically shows the parameter distribution in the residual space obtained by dividing the fine-tuning method of a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure.
[0064] like Figure 3B As shown in the figure, the horizontal and vertical coordinates represent the new feature dimensions obtained after the bit data is reduced in dimension. The markers of different shapes in the coordinates correspond to different features in the residual space. Figure 3B It can be seen that the distribution of marked points in the figure is relatively uniform, that is, the feature distribution is relatively scattered and lacks focus, which is more suitable for controlling the incremental function of learning objectives of pre-trained large models.
[0065] According to an embodiment of the present disclosure, the fine-tuning method for large models in computer hardware operation and maintenance also includes: determining a first weight matrix based on a key source basis, a key target basis and a first training matrix, wherein the first training matrix is determined based on a first low-rank matrix and a second low-rank matrix, the first low-rank matrix and the second low-rank matrix are obtained by random initialization, the number of rows of the first low-rank matrix is determined according to the number of columns of the key source basis, the number of columns of the second low-rank matrix is determined according to the number of rows of the key target basis, and the number of columns of the first low-rank matrix is the same as the number of rows of the second low-rank matrix; and determining a second weight matrix based on a supplementary source basis, a supplementary target basis and the second training matrix, wherein the second training matrix is determined based on a third low-rank matrix and a fourth low-rank matrix, the third low-rank matrix and the fourth low-rank matrix are obtained by random initialization, the number of rows of the third low-rank matrix is determined according to the number of columns of the supplementary source basis, the number of columns of the fourth low-rank matrix is determined according to the number of rows of the supplementary target basis, and the number of columns of the third low-rank matrix is the same as the number of rows of the fourth low-rank matrix.
[0066] According to the embodiments of the present disclosure, a primary space can be formed using the key source basis and the key target basis to maintain the original function. However, since the primary space is a high-dimensional space, the dimensions of the key source basis and the key target basis are also high. When adjusting parameters based on the key source basis and the key target basis, a large number of parameters need to be adjusted. Therefore, two low-rank matrices can be set based on the key source basis and the key target basis. By adjusting the parameters in the low-rank matrices, the matrix parameters corresponding to the entire primary space can be adjusted.
[0067] According to an embodiment of the present disclosure, a first low-rank matrix and a second low-rank matrix can be set. According to the properties of matrix multiplication, it can be determined that the number of rows of the first low-rank matrix is the same as the number of rows of the key source basis, the number of columns of the second low-rank matrix is the same as the number of columns of the key target basis, the number of columns of the first low-rank matrix is the same as the number of rows of the second low-rank matrix, and the number of columns of the first low-rank matrix is much smaller than the number of rows of the key source basis and the number of columns of the key target basis.
[0068] Among them, based on the key source basis, the key target basis and the first training matrix, the first weight matrix is determined, including: based on the key source basis and the first low-rank matrix, the feature direction in the main space is determined; based on the key target basis and the second low-rank matrix, the feature quantity in the main space is determined; and based on the feature direction in the main space and the feature quantity in the main space, the first weight matrix is determined.
[0069] According to an embodiment of the present disclosure, according to the key source substrate U p and the first low-rank matrix M p , which can determine the characteristic direction A in the main space p , that is, A p =U p M p , according to the key target base V p and the second low-rank matrix N p , can determine the feature quantity B in the main space p , that is, B p =V p N p .
[0070] According to the embodiment of the present disclosure, after determining the feature direction and feature quantity in the main space, the first weight matrix in the main space can be determined according to formula (2): :
[0071] (2)
[0072] According to formula (2), by adjusting M p With N p The parameters in can be used to realize the whole first weight matrix △W p Adjustment of parameters in . And because M p With N p Both are low-rank matrices, M p With N p The number of parameters in is much smaller than that of U p and V p Therefore, the adjustment of the first weight matrix in the main space can be achieved by adjusting a small number of parameters.
[0073] Similarly, based on the supplementary source basis, the supplementary target basis and the second training matrix, the second weight matrix is determined, including: determining the feature direction in the residual space based on the supplementary source basis and the third low-rank matrix; determining the feature quantity in the residual space based on the supplementary target basis and the fourth low-rank matrix; and determining the second weight matrix based on the feature direction in the residual space and the feature quantity in the residual space.
[0074] According to an embodiment of the present disclosure, according to the supplementary source substrate U res and the third low-rank matrix Mres , which can determine the feature direction A in the residual space res , that is, A res =U res M res , according to the key target base V res and the second low-rank matrix N res , which can determine the feature quantity B in the residual space res , that is, B res =V res N res .
[0075] According to the embodiment of the present disclosure, after determining the feature direction and feature quantity in the residual space, the second weight matrix in the residual space can be determined according to formula (3): :
[0076] (3)
[0077] According to formula (3), by adjusting M res With N res The parameters in can be used to realize the whole second weight matrix △W res Adjustment of parameters in . And because M res With N res Both are low-rank matrices, M res With N res The number of parameters in is much smaller than that of U res and V res Therefore, the adjustment of the second weight matrix in the main space can be achieved by adjusting a small number of parameters.
[0078] According to an embodiment of the present disclosure, the first training matrix is composed of a first low-rank matrix and a second low-rank matrix. The characteristic direction of the main space is determined according to the first low-rank matrix and the key source basis. The characteristic quantity of the main space is determined according to the second low-rank matrix and the key target basis. Further, the characteristics of the main space can be determined according to the first weight matrix composed of the characteristic direction and the characteristic quantity. Since the key source basis and the key target basis are both high-dimensional matrices, the number of parameters that need to be adjusted when updating the first weight matrix through the key source basis and the key target basis is large, and the parameter adjustment efficiency is low. However, by setting the first low-rank matrix and the second low-rank matrix and utilizing the properties of matrix multiplication, it is possible to achieve efficient updating of the first weight matrix by adjusting a small number of parameters in the first low-rank matrix and the second low-rank matrix, thereby improving the parameter adjustment efficiency.
[0079] According to an embodiment of the present disclosure, the weight matrix of the pre-trained large model includes an adjustable weight matrix and a fixed weight matrix; according to the target incremental function, the weight values in the first weight matrix and the second weight matrix are adjusted to obtain the adjusted first weight matrix and the adjusted second weight matrix, including: initializing the second low-rank matrix to a preset initial value so that the feature quantity in the initialized main space obtained after initialization meets the preset conditions; determining the initialization result of the first weight matrix based on the feature quantity in the initialized main space and the feature direction in the main space; processing the data input to the pre-trained large model based on the fixed weight matrix and the initialized first weight matrix to obtain output data; determining the model loss based on the output data and the expected output data; and adjusting the parameters of the first low-rank matrix and the second low-rank matrix to achieve parameter adjustment of the first weight matrix until the model loss converges.
[0080] According to an embodiment of the present disclosure, since training can be completed by only fine-tuning the model parameters during the incremental learning process, the parameter matrix learned by the pre-trained large model can be retained and used as a fixed weight matrix during the incremental learning process, wherein the fixed weight matrix does not need to be adjusted during the incremental learning process and can be used as the initial parameters of the pre-trained large model.
[0081] According to an embodiment of the present disclosure, at the beginning of incremental learning, the second low-rank matrix can be initialized to a preset initial value so that the feature quantity in the initialized main space obtained after initialization meets a preset condition. The preferred preset condition is that the feature quantity in the initialized main space obtained after initialization is the feature quantity in the fixed weight matrix, and the preset initial value is 0.
[0082] According to an embodiment of the present disclosure, the initialization result of the first weight matrix can be obtained by initializing the product between the feature quantity in the main space and the feature direction in the main space.
[0083] Similarly, the fourth low-rank matrix is initialized to a preset initial value so that the feature quantity in the initial residual space obtained after initialization satisfies the preset conditions. The initialization result of the second weight matrix can be obtained based on the product between the feature quantity in the initial residual space and the feature direction in the residual space.
[0084] According to an embodiment of the present disclosure, after determining the initialization result of the first weight matrix and the initialization result of the second weight matrix, the weight matrix of the pre-trained large model can be updated based on the fixed weight matrix, the initialized first weight matrix, and the initialized second weight matrix. The updated pre-trained large model is used to process the data input to the pre-trained large model to obtain output data corresponding to the input data. Based on the expected output data corresponding to the input data and the output data, the model loss can be determined, where the expected output data is the label of the input data.
[0085] According to an embodiment of the present disclosure, based on the model loss, the parameters of the first low-rank matrix and the second low-rank matrix are adjusted to complete the parameter adjustment of the first weight matrix, and the parameters of the third low-rank matrix and the fourth low-rank matrix are adjusted to complete the parameter adjustment of the second weight matrix. The above training process is repeated until the model loss converges, and the trained large model is determined based on the current weight matrix.
[0086] According to an embodiment of the present disclosure, after the model convergence is determined, the current weight matrix W' can be determined based on the first weight matrix after parameter adjustment, the second weight matrix after parameter adjustment, and the fixed weight matrix, as shown in formula (4):
[0087] (4)
[0088] in, is a fixed weight matrix, is the first weight matrix after parameter adjustment, is the second weight matrix after parameter adjustment.
[0089] Figure 4 The figure schematically shows a flowchart of initializing the weight matrix of a pre-trained large model according to the fine-tuning method of a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure.
[0090] like Figure 4 As shown, the weight matrix of the pre-trained large model includes a fixed weight matrix and an adjustable weight matrix. The adjustable weight matrix includes a first weight matrix and a second weight matrix. When initializing the first weight matrix, let B p =0, the initialization result of the first weight matrix can be made 0. Since the key target basis is fixed during the training process, it is only necessary to initialize the second low-rank matrix to 0. Similarly, initializing the fourth low-rank matrix to 0 can make B res =0, so that the initialization result of the second weight matrix is 0. After completing the initialization of the second low-rank matrix and the fourth low-rank matrix, the initialized weight matrix can be obtained according to formula (4). Since the first weight matrix and the second weight matrix in the adjustable weight matrix are both 0, the initialized weight matrix is actually the same as the fixed weight matrix, which can ensure that the subsequent fine-tuning training process is carried out on the basis of the fixed weight matrix, thereby avoiding the impact of initialization on the model parameters.
[0091] According to an embodiment of the present disclosure, the fine-tuning method of a large model in computer hardware operation and maintenance also includes: determining a first learning rate of a first weight matrix based on a preset learning rate reference value and an eigenvalue of a main space; determining a second learning rate of a second weight matrix based on a preset learning rate reference value and an eigenvalue of a residual space; and training the first training matrix using the first learning rate and training the second training matrix using the second learning rate.
[0092] According to the embodiments of the present disclosure, the Hessian eigenvalue distributions of the main space and the residual space are usually different. Often much larger than the eigenvalue of the residual space In this case, if a global unified learning rate is used during the training of the pre-trained large model , which will result in the main space and residual space being unable to be trained efficiently at the same time.
[0093] Specifically, if , the main space can converge quickly, and the convergence factor of the residual space is ,according to and The relationship between the size of the residual space shows that the convergence factor is close to 1, and the convergence speed is very slow. Correspondingly, if we take , so that the residual space can converge quickly, then the convergence factor of the main space It will be greater than 2, causing the main space to diverge and fail to converge.
[0094] According to an embodiment of the present disclosure, independent learning rates can be set for the main space and the residual space respectively. The first learning rate and the second learning rate , making and Satisfy the relationship shown in formula (5):
[0095] (5)
[0096] According to an embodiment of the present disclosure, according to formula (5), a preset learning rate reference value a can be set, and formula (6) is obtained:
[0097] (6)
[0098] Therefore, according to the preset learning rate reference value a and the eigenvalue of the main space , the first learning rate of the first weight matrix can be determined Similarly, according to the preset learning rate reference value a and the eigenvalue of the residual space , the second learning rate of the second weight matrix can be determined .
[0099] According to an embodiment of the present disclosure, the first learning rate and the second learning rate are calculated and the proportional relationship between the two is controlled based on the eigenvalues of the main space and the eigenvalues of the residual space, so that the training efficiency and accuracy in the main space and the residual space are balanced, thereby improving the overall parameter adjustment efficiency of the pre-trained large model. By setting a preset learning rate reference value, the preset learning rate reference value can be changed to change the values of the first learning rate and the second learning rate while ensuring that the relative ratio of the first learning rate and the second learning rate remains unchanged, thereby changing the parameter adjustment speed of the pre-trained large model.
[0100] According to an embodiment of the present disclosure, the fine-tuning method for large models in computer hardware operation and maintenance also includes: determining a first parameter amount to be adjusted in the main space based on a first training matrix; determining a second parameter amount to be adjusted in the residual space based on a second training matrix; adjusting the first norm ratio based on the first parameter amount to be adjusted and the second parameter amount to be adjusted to obtain an adjusted first norm ratio; and regularizing the main space and the residual space respectively based on the first parameter amount to be adjusted, the second parameter amount to be adjusted, the adjusted first norm ratio and the second norm ratio.
[0101] According to the embodiment of the present disclosure, the first low-rank matrix and the second low-rank matrix constituting the first training matrix are concatenated to determine the total amount of parameters that need to be adjusted in the main space, that is, the first parameter amount to be adjusted By concatenating the third low-rank matrix and the fourth low-rank matrix that constitute the second training matrix, the total amount of parameters that need to be adjusted in the residual space can be determined, that is, the second amount of parameters to be adjusted .
[0102] According to the embodiments of the present disclosure, during each iterative parameter adjustment process, the ratio of the actual adjusted parameter amount to the parameter amount to be adjusted in the main space and the residual space can be controlled by setting the norm ratio. Specifically, a first norm ratio corresponding to the main space and a second norm ratio corresponding to the residual space can be set, and the norm ratio penalty term can be calculated using formula (7): :
[0103] (7)
[0104] Among them, γ p is the first norm ratio, γ res is the second norm ratio, ||·|| F is the F-norm, is the square value of the F-norm, μ and ϵ are hyperparameters, μ is the penalty coefficient, and ϵ is a small constant used to prevent The denominator of the term is 0.
[0105] According to formula (7), the norm ratio penalty term can be calculated by the first norm ratio and the second norm ratio. When the norm ratio penalty term does not converge, the first norm ratio and the second norm ratio are adjusted. The norm ratio penalty term can be used to constrain model parameters to prevent overfitting.
[0106] According to an embodiment of the present disclosure, when it is determined by formula (7) that the norm proportional penalty term has not converged, γ can be calculated by formula (8): p Make iterative adjustments:
[0107] (8)
[0108] Among them, t is the current iteration round, α is the hyperparameter, γ p The minimum value of γ p The maximum value of the function, clamp(compute, min, max) is a limiting function used to control the output range of the function. For example, when the calculation result of compute is less than min, the function output is min; when the calculation result of compute is greater than max, the function output is max; when the calculation result of compute is within the interval [min, max], the function output is the calculation result of compute.
[0109] According to an embodiment of the present disclosure, the first parameter to be adjusted θ in the main space p The second parameter to be adjusted relative to the residual space θ res When the value is large, the penalty term in formula (7) Large, using formula (8), it is possible to p Relative θ res In the case of large p , and then γ p Substituting into formula (7), p Apply a stronger constraint so that θ p Relative θ res Reduce, so that the penalty term is reduced and the control norm proportional penalty term converges.
[0110] According to embodiments of the present disclosure, when convergence of the norm ratio penalty term is determined, the main space can be regularized based on the adjusted first norm ratio, and the residual space can be regularized based on the second norm ratio. This allows for faster convergence of the main and residual spaces, more stable training, and improved training effectiveness and efficiency.
[0111] According to the embodiments of the present disclosure, by setting and calculating the norm ratio penalty term, the parameter amounts that need to be adjusted in the main space and the residual space can be obtained. When the norm ratio penalty term diverges, the first norm ratio is iteratively adjusted. The parameter amounts that need to be adjusted in the main space can be controlled through the first norm ratio, thereby adjusting the ratio between the parameter amounts that need to be adjusted in the main space and the parameter amounts that need to be adjusted in the residual space to stabilize the core adjustment of the main space, so that the pre-trained large model will not lose existing functions during fine-tuning training, and the residual space can flexibly learn task details, so that the residual space has higher learning efficiency for the target incremental function.
[0112] Figure 5 A block diagram of a device for fine-tuning a large model in computer hardware operation and maintenance according to an embodiment of the present disclosure is schematically shown.
[0113] like Figure 5 As shown, the fine-tuning device 500 for a large model in computer hardware operation and maintenance includes a matrix determination module 510 and a matrix adjustment module 520.
[0114] Specifically, the matrix determination module 510 is used to determine the adjustable weight matrix that needs to be fine-tuned for the pre-trained large model. The adjustable weight matrix includes a first weight matrix corresponding to the main space of the feature space of the pre-trained large model, and a second weight matrix corresponding to the residual space of the feature space of the pre-trained large model. The main space represents the feature space where the functional characteristics of the original function of the pre-trained large model are located, and the residual space represents the feature space where the functional characteristics of the target incremental function are located. The original function includes the recommendation function and pattern recognition function corresponding to the computer hardware, and the target incremental function includes the decision function corresponding to the computer hardware. The matrix parameters in the main space and the residual space correspond to the computer hardware parameters, the parameters in the feature space of the original function correspond to the recommendation function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function correspond to the decision function parameters of the computer hardware.
[0115] The matrix adjustment module 520 is used to adjust the weight values in the first weight matrix and the second weight matrix according to the target incremental function to obtain an adjusted first weight matrix and an adjusted second weight matrix, wherein the adjustment of the second weight matrix is performed based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is performed based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space. The adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
[0116] According to an embodiment of the present disclosure, the fine-tuning device 500 for a large model in computer hardware operation and maintenance further includes a matrix decomposition module, a first basis determination module, a second basis determination module, and a space determination module.
[0117] A matrix decomposition module is used to perform matrix decomposition on the adjustable weight matrix to determine the source basis of the adjustable weight matrix in the source space and the target basis in the target space, wherein the source basis is used to map the data input to the pre-trained large model from the source feature space to the target feature space, and the target basis is used to map the data that has completed data transformation in the target space back to the source feature space. The source basis includes multiple source vectors, and the target basis includes multiple target vectors.
[0118] The first basis determination module is used to form a key source basis from a preset number of source vectors in the source basis according to the amount of feature information included in the source vector, and to form a supplementary source basis from other source vectors.
[0119] The second basis determination module is used to form a key target basis from a preset number of target vectors in the target basis and to form a supplementary target basis from other target vectors according to the amount of feature information included in the target vector.
[0120] The space determination module is used to determine the main space based on the key source basis and the key target basis, and to determine the residual space based on the supplementary source basis and the supplementary target basis.
[0121] According to an embodiment of the present disclosure, the fine-tuning device 500 for a large model in computer hardware operation and maintenance further includes a first matrix determination module and a second matrix determination module.
[0122] The first matrix determination module is used to determine the first weight matrix based on the key source basis, the key target basis and the first training matrix, wherein the first training matrix is determined according to the first low-rank matrix and the second low-rank matrix, the first low-rank matrix and the second low-rank matrix are obtained by random initialization, the number of rows of the first low-rank matrix is determined according to the number of columns of the key source basis, the number of columns of the second low-rank matrix is determined according to the number of rows of the key target basis, and the number of columns of the first low-rank matrix is the same as the number of rows of the second low-rank matrix.
[0123] The second matrix determination module is used to determine the second weight matrix based on the supplementary source basis, the supplementary target basis and the second training matrix, wherein the second training matrix is determined based on the third low-rank matrix and the fourth low-rank matrix, the third low-rank matrix and the fourth low-rank matrix are obtained by random initialization, the number of rows of the third low-rank matrix is determined based on the number of columns of the supplementary source basis, the number of columns of the fourth low-rank matrix is determined based on the number of rows of the supplementary target basis, and the number of columns of the third low-rank matrix is the same as the number of rows of the fourth low-rank matrix.
[0124] According to an embodiment of the present disclosure, the first matrix determination module includes a direction determination submodule, a feature quantity determination submodule, and a matrix determination submodule.
[0125] The direction determination submodule is used to determine the feature direction in the main space based on the key source basis and the first low-rank matrix.
[0126] The feature quantity determination submodule is used to determine the feature quantity in the main space based on the key target basis and the second low-rank matrix.
[0127] The matrix determination submodule is used to determine a first weight matrix based on the feature direction in the main space and the feature quantity in the main space.
[0128] According to an embodiment of the present disclosure, the matrix adjustment module 520 includes a matrix initialization submodule, an initialization result determination submodule, a result output submodule, a loss determination submodule, and a parameter adjustment submodule.
[0129] The matrix initialization submodule is used to initialize the second low-rank matrix to a preset initial value so that the feature quantity in the initialized main space obtained after initialization meets the preset conditions.
[0130] The initialization result determination submodule is used to determine the initialization result of the first weight matrix based on the feature quantity and the feature direction in the initialized main space.
[0131] The result output submodule is used to process the data input to the pre-trained large model based on the fixed weight matrix and the initialized first weight matrix to obtain output data.
[0132] The loss determination submodule is used to determine the model loss based on the output data and the expected output data.
[0133] The parameter adjustment submodule is used to adjust the parameters of the first low-rank matrix and the second low-rank matrix to achieve parameter adjustment of the first weight matrix until convergence is achieved according to the model loss.
[0134] According to an embodiment of the present disclosure, the fine-tuning device 500 for a large model in computer hardware operation and maintenance further includes a first learning rate determination module, a second learning rate determination module, and a matrix training module.
[0135] The first learning rate determination module is used to determine a first learning rate of the first weight matrix according to a preset learning rate reference value and an eigenvalue of the main space.
[0136] The second learning rate determination module is used to determine the second learning rate of the second weight matrix according to a preset learning rate reference value and an eigenvalue of the residual space.
[0137] The matrix training module is used to train the first training matrix using a first learning rate and to train the second training matrix using a second learning rate.
[0138] According to an embodiment of the present disclosure, the fine-tuning device 500 for a large model in computer hardware operation and maintenance further includes a first parameter quantity determination module, a second parameter quantity determination module, a scale adjustment module, and a spatial regularization module.
[0139] The first parameter determination module is used to determine a first parameter to be adjusted in the main space based on the first training matrix.
[0140] The second parameter determination module is used to determine the second parameter to be adjusted in the residual space based on the second training matrix.
[0141] The ratio adjustment module is used to adjust the first norm ratio based on the first parameter amount to be adjusted and the second parameter amount to be adjusted to obtain an adjusted first norm ratio.
[0142] The spatial regularization module is used to regularize the main space and the residual space respectively based on the first parameter to be adjusted, the second parameter to be adjusted, the adjusted first norm ratio and the adjusted second norm ratio.
[0143] It is understood that the matrix determination module 510 and the matrix adjustment module 520 can be implemented in a single module, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the matrix determination module 510 and the matrix adjustment module 520 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented in hardware or firmware using any other reasonable method of circuit integration or packaging, or can be implemented using a suitable combination of software, hardware, and firmware. Alternatively, at least one of the matrix determination module 510 and the matrix adjustment module 520 can be at least partially implemented as a computer program module, which, when executed by a computer, can perform the functions of the corresponding module.
[0144] Figure 6A schematic block diagram of an example electronic device that can be used to implement the method of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0145] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0146] Multiple components in the electronic device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0147] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the application execution method. For example, in some embodiments, the application execution method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the application execution method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the application execution method via any other suitable means (e.g., via firmware).
[0148] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0152] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0153] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem, addressing the management difficulties and limited scalability of traditional physical hosts and VPS ("Virtual Private Server") services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0154] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0155] Although the present disclosure has been shown and described with reference to certain exemplary embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made to the present disclosure without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure should not be limited to the above-described embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims.
Claims
1. A method for fine-tuning a large model in computer hardware operation and maintenance, comprising: Determine an adjustable weight matrix that needs to be fine-tuned for a pre-trained large model, the adjustable weight matrix including a first weight matrix corresponding to a main space of a feature space of the pre-trained large model, and a second weight matrix corresponding to a residual space of the feature space of the pre-trained large model, the main space representing the feature space in which the functional characteristics of the original function of the pre-trained large model are located, the residual space representing the feature space in which the functional characteristics of the target incremental function are located, the original function including a recommendation function and a pattern recognition function corresponding to the computer hardware, the target incremental function including a decision function corresponding to the computer hardware, the matrix parameters in the main space and the residual space corresponding to computer hardware parameters, the parameters in the feature space of the original function corresponding to the recommendation function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function corresponding to the decision function parameters of the computer hardware; as well as According to the target incremental function, the weight values in the first weight matrix and the second weight matrix are adjusted to obtain an adjusted first weight matrix and an adjusted second weight matrix, wherein the adjustment of the second weight matrix is performed based on the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is performed based on the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space, the adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
2. The method according to claim 1, further comprising: Performing matrix decomposition on the adjustable weight matrix to determine a source basis of the adjustable weight matrix in a source space and a target basis of the adjustable weight matrix in a target space, wherein the source basis is used to map data input into the pre-trained large model from a source feature space to a target feature space, and the target basis is used to map data that has completed data transformation in the target space back to the source feature space, the source basis includes a plurality of source vectors, and the target basis includes a plurality of target vectors; According to the amount of feature information included in the source vectors, a preset number of source vectors in the source basis form a key source basis, and other source vectors form a supplementary source basis; According to the amount of feature information included in the target vector, the preset number of target vectors in the target basis constitute a key target basis, and the other target vectors constitute a supplementary target basis; The primary space is determined based on the key source basis and the key target basis, and the residual space is determined based on the supplementary source basis and the supplementary target basis.
3. The method according to claim 2, further comprising: determining the first weight matrix based on the key source basis, the key target basis, and a first training matrix, wherein the first training matrix is determined according to a first low-rank matrix and a second low-rank matrix, the first low-rank matrix and the second low-rank matrix are obtained by random initialization, the number of rows of the first low-rank matrix is determined according to the number of columns of the key source basis, the number of columns of the second low-rank matrix is determined according to the number of rows of the key target basis, and the number of columns of the first low-rank matrix is the same as the number of rows of the second low-rank matrix; and Based on the supplementary source basis, the supplementary target basis and the second training matrix, the second weight matrix is determined, wherein the second training matrix is determined according to the third low-rank matrix and the fourth low-rank matrix, the third low-rank matrix and the fourth low-rank matrix are obtained by random initialization, the number of rows of the third low-rank matrix is determined according to the number of columns of the supplementary source basis, the number of columns of the fourth low-rank matrix is determined according to the number of rows of the supplementary target basis, and the number of columns of the third low-rank matrix is the same as the number of rows of the fourth low-rank matrix.
4. The method according to claim 3, wherein: The determining the first weight matrix based on the key source basis, the key target basis, and the first training matrix includes: Determining characteristic directions in a primary space based on the key source basis and the first low-rank matrix; Determining a feature quantity in a main space based on the key target basis and the second low-rank matrix; and The first weight matrix is determined based on the feature direction in the main space and the feature quantity in the main space.
5. The method according to claim 4, wherein the weight matrix of the pre-trained large model includes the adjustable weight matrix and the fixed weight matrix; The step of adjusting the weight values in the first weight matrix and the second weight matrix according to the target increment function to obtain the adjusted first weight matrix and the adjusted second weight matrix includes: Initializing the second low-rank matrix to a preset initial value so that the feature quantity in the initialized main space obtained after initialization meets a preset condition; Determining an initialization result of the first weight matrix based on the feature quantity in the initialized main space and the feature direction in the main space; Processing the data input to the pre-trained large model based on the fixed weight matrix and the initialized first weight matrix to obtain output data; determining a model loss based on the output data and the expected output data; as well as Parameters of the first low-rank matrix and the second low-rank matrix are adjusted to achieve parameter adjustment of the first weight matrix until convergence according to the model loss.
6. The method according to claim 3, further comprising: Determining a first learning rate of the first weight matrix according to a preset learning rate reference value and an eigenvalue of the main space; Determining a second learning rate of the second weight matrix according to the preset learning rate reference value and the eigenvalue of the residual space; as well as The first training matrix is trained using the first learning rate, and the second training matrix is trained using the second learning rate.
7. The method according to claim 6, further comprising: Determining a first parameter to be adjusted in the main space based on the first training matrix; Determining a second parameter to be adjusted in the residual space based on the second training matrix; Adjusting the first norm ratio based on the first parameter to be adjusted and the second parameter to be adjusted to obtain an adjusted first norm ratio; and Regularization is performed on the main space and the residual space based on the first parameter amount to be adjusted, the second parameter amount to be adjusted, the adjusted first norm ratio, and the adjusted second norm ratio.
8. A fine-tuning device for large models in computer hardware operation and maintenance, comprising: a matrix determination module, for determining an adjustable weight matrix that needs to be fine-tuned for a pre-trained large model, the adjustable weight matrix comprising a first weight matrix corresponding to a main space of a feature space of the pre-trained large model, and a second weight matrix corresponding to a residual space of the feature space of the pre-trained large model, the main space representing the feature space in which the functional characteristics of the original function of the pre-trained large model are located, the residual space representing the feature space in which the functional characteristics of the target incremental function are located, the matrix parameters in the main space and the residual space corresponding to computer hardware parameters, the parameters in the feature space of the original function corresponding to the recommended function and pattern recognition function parameters of the computer hardware, and the parameters in the feature space of the target incremental function corresponding to the decision-making function parameters of the computer hardware; as well as A matrix adjustment module is used to adjust the weight values in the first weight matrix and the second weight matrix according to the target incremental function to obtain an adjusted first weight matrix and an adjusted second weight matrix, wherein the adjustment of the second weight matrix is performed according to the decision function parameters corresponding to the computer hardware, and the adjustment of the first weight matrix is performed according to the recommended function and pattern recognition function parameters corresponding to the computer hardware in the main space, the adjusted first weight matrix is used to represent the functional characteristics of the original function, and the adjusted second weight matrix represents the functional characteristics of the target incremental function.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to claims 1 to 7.