LoRA weight fusion method and device for large language model inference system

By obtaining multiple LoRA weight data in a large language model inference system, determining their fusion ratio and performing splicing or segmentation processing, LoRA fusion weights are generated, the problem of dynamic multi-style fusion is solved and the flexibility and performance of the system is improved.

CN120354930AActive Publication Date: 2025-07-22BEIJING SILICONFLOW TECHNOLOGY CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510077239.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-07-22
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing large language model inference system is difficult to dynamically adjust the style proportion of multiple LoRA models or generate text of multiple styles at the same time, which limits its application potential in multi-style text generation.

Method used

By obtaining multiple LoRA weight data, determining their fusion ratios, and splicing or slicing based on these ratios, LoRA fusion weights are generated, and these weights are used for inference calculations to achieve dynamic multi-style fusion.

Benefits of technology

It improves the memory utilization rate, meets the dual requirements of flexibility and performance of large language model inference systems, and realizes the flexibility and efficient calculation of multi-style text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354930A_ABST
    Figure CN120354930A_ABST
Patent Text Reader

Abstract

The invention relates to a LoRA weight fusion method and device for a large language model inference system. The method comprises the steps that a large language model inference system obtains multiple pieces of LoRA weight data; determining a fusion proportion corresponding to the multiple pieces of LoRA weight data; carrying out splicing processing or segmentation processing on the multiple pieces of LoRA weight data based on the fusion proportion to generate a LoRA fusion weight; the inference system of the large language model obtains input data; and calling the LoRA fusion weight based on the input data to perform reasoning calculation. The LoRA weight fusion method and device for the large language model inference system are suitable for scenes needing dynamic multi-style fusion, the video memory utilization rate can be improved, and the dual requirements of the large language model inference system for flexibility and performance are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer information processing. Specifically, it relates to a LoRA weight fusion method and device for a large language model inference system. Background Art

[0002] In recent years, deep learning technology has made significant progress in fields such as natural language processing (NLP), computer vision (CV), and reinforcement learning (RL). Among them, large language models (LLMs), as a type of deep learning model, are mainly used to process text data. These models learn rich language knowledge and semantic information through pre-training on large-scale corpora and are widely applied to various downstream tasks such as text generation, machine translation, and question-answering systems.

[0003] Although large language models already have strong language understanding capabilities in the pre-training stage, for the actual needs of specific tasks, further performance improvement is still required through fine-tuning. Fine-tuning refers to optimizing and training a pre-trained model using a small amount of labeled data for a specific task, enabling the model to better adapt to the data characteristics and requirements of the specific task, thereby achieving better actual performance.

[0004] To reduce the computational cost of fine-tuning and improve efficiency, an efficient model fine-tuning method, namely LoRA (Low-Rank Adaptation), has been proposed in recent years. The core idea of LoRA is to add low-rank matrices on the basis of a pre-trained model to achieve fine-tuning, and only a small number of parameters need to be adjusted to significantly improve the model performance.

[0005] However, in practical applications, the style fusion of multiple LoRA models has become a new challenge. For example, LoRA models trained on datasets with different styles can enable the model to generate text in a specific style, but existing methods are difficult to dynamically adjust the style proportion or generate text in multiple styles simultaneously. The current LLM inference system does not support the function of dynamically fusing multiple LoRA models, which limits its application potential in multi-style text generation.

[0006] Therefore, a new LoRA weight fusion method and device for a large language model inference system are needed.

[0007] The above information disclosed in the background art section is only used to enhance the understanding of the background of this application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0008] In view of this, the present application provides a LoRA weight fusion method and device for a large language model inference system, which is applicable to scenarios that require dynamic multi-style fusion, can improve video memory utilization, and meet the dual requirements of flexibility and performance of the large language model inference system.

[0009] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part through the practice of the present application.

[0010] According to one aspect of the present application, a LoRA weight fusion method for a large language model inference system is proposed. The method includes: the large language model inference system obtains a plurality of LoRA weight data; determines the fusion ratio corresponding to the plurality of LoRA weight data; based on the fusion ratio, performs splicing processing or splitting processing on the plurality of LoRA weight data to generate LoRA fusion weights; the inference system of the large language model obtains input data; and performs inference calculation by invoking the LoRA fusion weights based on the input data.

[0011] In an exemplary embodiment of the present application, determining the fusion ratio corresponding to the plurality of LoRA weight data includes: determining the fusion ratio corresponding to the plurality of LoRA weight data according to preset task requirements.

[0012] In an exemplary embodiment of the present application, performing splicing processing or splitting processing on the plurality of LoRA weight data based on the fusion ratio to generate LoRA fusion weights includes: performing splicing processing or splitting processing on the plurality of LoRA weight data based on the rank dimension of the LoRA weight data; and fusing the data after the splicing processing or splitting processing according to the fusion ratio to generate the LoRA fusion weights.

[0013] In an exemplary embodiment of the present application, performing splicing processing or splitting processing on the plurality of LoRA weight data based on the rank dimension of the LoRA weight data to generate the LoRA processing weights includes: when the rank dimension of the LoRA weight data is less than the dimension threshold, performing splicing processing on the plurality of LoRA weight data and summarizing them into the LoRA processing weights.

[0014] In an exemplary embodiment of the present application, fusing the data after the splicing processing or splitting processing according to the fusion ratio to generate the LoRA fusion weights includes: according to the fusion ratio, performing matrix splicing on the LoRA processing weights according to the rank dimension to generate the LoRA fusion weights.

[0015] In an exemplary embodiment of the present application, splicing or splitting the multiple LoRA weight data based on the rank dimension of the LoRA weight data to generate the LoRA processed weight includes: when the rank dimension of the LoRA weight data is greater than the dimension threshold, splitting the LoRA weight data into multiple LoRA processed weights.

[0016] In an exemplary embodiment of the present application, fusing the data after splicing or splitting according to the fusion ratio to generate the LoRA fused weight further includes: after performing the splitting process, storing multiple LoRA processed weights and their corresponding fusion ratios based on the paging management method to generate the LoRA fused weight.

[0017] In an exemplary embodiment of the present application, invoking the LoRA fused weight for inference calculation based on the input data includes: performing inference calculation on the input data and the fused weight to generate a LoRA result; adding the LoRA result and the output result of the large language model to generate an inference result.

[0018] In an exemplary embodiment of the present application, performing inference calculation on the input data and the fused weight to generate a LoRA result includes: after performing the splitting process, performing inference calculation on the input data with multiple LoRA processed weights and their corresponding fusion ratios respectively to generate a LoRA result.

[0019] According to one aspect of the present application, a LoRA weight fusion device for a large language model inference system is proposed. The device includes: a data module for obtaining multiple LoRA weight data for the large language model inference system; a ratio module for determining the fusion ratios corresponding to the multiple LoRA weight data; a processing module for splicing or splitting the multiple LoRA weight data based on the fusion ratios to generate a LoRA fused weight; an input module for obtaining input data for the inference system of the large language model; and an inference module for invoking the LoRA fused weight for inference calculation based on the input data.

[0020] According to one aspect of the present application, an electronic device is proposed. The electronic device includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method as described above.

[0021] According to one aspect of the present application, a computer-readable medium is proposed, on which a computer program is stored. When the program is executed by a processor, it implements the method as described above.

[0022] The LoRA weight fusion method and device for the large language model inference system according to the present application obtain multiple LoRA weight data through the large language model inference system; determine the fusion ratios corresponding to the multiple LoRA weight data; splice or segment the multiple LoRA weight data based on the fusion ratios to generate LoRA fusion weights; the inference system of the large language model obtains input data; and performs inference calculations by calling the LoRA fusion weights based on the input data. This method is applicable to scenarios that require dynamic multi-style fusion, can improve the video memory utilization rate, and meet the dual requirements of flexibility and performance of the large language model inference system.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objectives, features, and advantages of the present application will become more apparent. The following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of a LoRA weight fusion method for a large language model inference system shown according to an exemplary embodiment.

[0026] Figure 2 is a flowchart of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0027] Figure 3 is a schematic diagram of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0028] Figure 4 is a schematic diagram of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0029] Figure 5 is a flowchart of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0030] Figure 6 is a schematic diagram of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0031] Figure 7Schematic diagram of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment.

[0032] Figure 8 Block diagram of a LoRA weight fusion device for a large language model inference system shown according to an exemplary embodiment.

[0033] Figure 9 Block diagram of an electronic device shown according to an exemplary embodiment.

[0034] Figure 10 Block diagram of a computer-readable medium shown according to an exemplary embodiment. Detailed implementation manners

[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0036] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.

[0037] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0039] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below may be referred to as the second component without departing from the teachings of the concepts of this application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0040] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily required for implementing this application, so they cannot be used to limit the protection scope of this application.

[0041] Figure 1 It is a flowchart of a LoRA weight fusion method for a large language model inference system shown according to an exemplary embodiment. The LoRA weight fusion method 10 for a large language model inference system at least includes steps S102 to S110.

[0042] As Figure 1 shown, in S102, the large language model inference system obtains a plurality of LoRA weight data.

[0043] In the actual application process, the large language model inference system obtains a plurality of LoRA weight data through loading or initialization operations. These weight data can be from LoRA models of different styles, such as LoRA model weights trained for specific data sets such as humorous style and ironic style respectively. Each LoRA weight data contains a low-rank matrix representation of the original model parameters and has independent style features.

[0044] In S104, the fusion ratios corresponding to the plurality of LoRA weight data are determined. For example, the fusion ratios corresponding to the plurality of LoRA weight data are determined according to preset task requirements.

[0045] In a specific embodiment, for example, the inference system determines the fusion ratios of the plurality of LoRA weight data according to preset task requirements. The fusion ratios can be dynamically calculated based on user input, task description, target style requirements, etc. For example, when generating text, 80% of the humorous style and 20% of the ironic style are required, then the corresponding weight ratios are assigned as 0.8 and 0.2.

[0046] For another example, the ratios can be stored as corresponding fusion coefficients for convenient direct calling in subsequent calculations.

[0047] In S106, based on the fusion ratios, the plurality of LoRA weight data are subjected to splicing processing or splitting processing to generate LoRA fusion weights.

[0048] It is worth mentioning that LoRA Linear adds LoRA-related calculations on top of the original Base Linear. The calculation process of Base Linear is as follows: base_output = input @ weight.T In the LoRA Linear structure, based on the original weight tensor weight, two additional tensors LoRA_A and LoRA_B are added. The calculation method is to calculate the multiplication of two small matrices additionally, and then add the output result of LoRA to the output result of BaseLinear to obtain the final output result: LoRA_output = input @ LoRA_A.T @ LoRA_B.T # (@ represents matrix multiplication,.T represents matrix transpose) output = base_output + LoRA_output Among them, the shapes of each tensor are as follows: plaintext input: (batch_size, in_features) weight: (out_features, in_features) base_output: (batch_size, out_features) LoRA_A: (rank, in_features) LoRA_B: (out_features, rank) LoRA_output: (batch_size, out_features) output: (batch_size, out_features) Because the shapes of base_output and LoRA_output are the same, they can be directly added together.

[0049] From this, it can be known that the LoRA weights have the property of being splittable in the rank dimension. Specifically, a LoRA model with a larger rank can be split into several smaller LoRA models in the rank dimension. By calculating the output results of these split LoRA models separately and adding them together, the final result is equivalent to the calculation result without splitting.

[0050] For example: LoRA_output_0 = input @ LoRA_A_0.T @ LoRA_B_0.T LoRA_output_1 = input @ LoRA_A_1.T @ LoRA_B_1.T LoRA_output = LoRA_output_0 + LoRA_output_1 Assume that a LoRA model is equally divided into two smaller models, then the shapes of each tensor are as follows: plaintext LoRA_A_0: (rank / 2, in_features) LoRA_B_0: (out_features, rank / 2) LoRA_A_1: (rank / 2, in_features) LoRA_B_1: (out_features, rank / 2) LoRA_output_0: (batch_size, out_features) LoRA_output_1: (batch_size, out_features) Since the rank dimension only appears in the intermediate results of LoRA calculations, splitting the rank dimension will not affect the final calculation results.

[0051] The splittability of LoR weights in the rank dimension is closely related to their concatenability. Specifically, if multiple small LoRA models need to be calculated simultaneously, they can be concatenated in the rank dimension into a larger LoRA model, and then the result derivation can be completed through a unified calculation process.

[0052] plaintext LoRA_A_combined = concat(LoRA_A_0, LoRA_A_1, dim = 0) # Concatenate along the rank dimension LoRA_B_combined = concat(LoRA_B_0, LoRA_B_1, dim = 1) # Concatenate along the rank dimension LoRA_output_combined = input @ LoRA_A_combined.T @ LoRA_B_combined.T Through splicing, the calculations of multiple LoRA models can be completed at once, reducing the complexity of calculating each individual model one by one while ensuring that the result is consistent with the result obtained by calculating each one separately and then adding them together.

[0053] Through the above characteristics, the flexibility of the LoRA model is greatly improved. It can either split out sub-models along the rank dimension when needed or splice multiple small models into a unified large model in a task scenario to achieve more efficient inference calculations.

[0054] In this application, for example, the multiple LoRA weight data can be spliced or split based on the rank dimension of the LoRA weight data; the data after splicing or splitting is fused according to the fusion ratio to generate the LoRA fused weight.

[0055] Splicing or splitting multiple LoRA weight data to generate a unified LoRA fused weight.

[0056] In one embodiment, the splicability of LoRA weights along the rank dimension can be utilized to splice the weights of different LoRA models along the rank dimension to form an overall weight matrix. During the splicing process, appropriate padding can be performed to ensure dimensional consistency after splicing to adapt to the rank sizes of different LoRA models.

[0057] In another embodiment, if the rank of a certain LoRA model is large, it can first be split into multiple sub-weight matrices with smaller ranks along the rank dimension. The split sub-weights can respectively participate in subsequent fusion calculations to improve the video memory utilization efficiency.

[0058] More specifically, the spliced or split LoRA weight matrix can be weighted and fused according to a preset fusion ratio. The weighting process can be completed through matrix multiplication operations to ensure that LoRA weights of different styles participate in the calculations according to the ratio.

[0059] In S108, the inference system of the large language model obtains the input data. The inference system of the large language model receives the input data from the user side, such as the starting text or context prompt in a text generation task. The input data will serve as the basis for inference calculations and directly affect the final output result.

[0060] In S110, the LoRA fused weight is called for inference calculations based on the input data. For example, the input data and the fused weight can be used for inference calculations to generate a LoRA result; the LoRA result and the output result of the large language model are added together to generate an inference result.

[0061] The LoRA weight fusion method for large language model inference systems according to this application obtains multiple LoRA weight data through the large language model inference system; determines the fusion ratios corresponding to the multiple LoRA weight data; performs splicing or splitting processing on the multiple LoRA weight data based on the fusion ratios to generate LoRA fusion weights; the inference system of the large language model obtains input data; and performs inference calculations by calling the LoRA fusion weights based on the input data. This method is applicable to scenarios that require dynamic multi-style fusion, can improve video memory utilization, and meet the dual requirements of flexibility and performance for large language model inference systems.

[0062] It should be clearly understood that this application describes how to form and use specific examples, but the principles of this application are not limited to any details of these examples. Instead, based on the teachings disclosed in this application, these principles can be applied to many other embodiments.

[0063] Figure 2 It is a flowchart of a LoRA weight fusion method for a large language model inference system shown according to another exemplary embodiment. Figure 2 The shown process 20 is a detailed description of Figure 1 the dimension splicing method in the shown process.

[0064] As Figure 2 shown, in S202, when the rank dimension of the LoRA weight data is less than the dimension threshold, splicing processing is performed on the multiple LoRA weight data. First, the rank dimension of the obtained LoRA weight data can be judged. If it is judged that the rank dimension of the weight data is less than the preset dimension threshold, splicing processing is performed on it.

[0065] In S204, according to the fusion ratio, the LoRA processed weights are matrix-spliced according to the rank dimension to generate the LoRA fusion weights.

[0066] In S206, inference calculations are performed on the input data and the fusion weights to generate LoRA results. Inference calculations are performed on the input data and the generated LoRA fusion weights to generate the calculation results of LoRA. In this process, the spliced fusion weights can effectively combine the characteristics of different LoRA models to achieve diverse inference outputs.

[0067] In S208, the LoRA results and the output results of the large language model are added together to generate inference results. The calculation results of LoRA are added to the original output results of the large language model to generate the final inference results. Through this weighted fusion method, the inference system can not only retain the general knowledge of the large language model but also improve the performance of specific tasks through the characteristics of the LoRA model.

[0068] In a specific embodiment, in combination with the splittability of LoRA weights in the rank dimension introduced above, multiple LoRAs can be spliced in the rank dimension and then calculated.

[0069] Figure 3 This is the storage format when a LoRA model with rank = 2 and a LoRA model with rank = 1 are fused. The yellow part represents the weights of LoRA_model_0 with rank = 2, and the pink part represents the weights of LoRA_model_1 with rank = 1. The white part is the unused video memory, and the gray part is padding. Further, a tensor with a shape of (max_LoRAs, rank) can be defined to represent the proportion of each LoRA model in the fusion process. Since it is spliced in the rank dimension, there are rank elements accordingly.

[0070] According to the implementation manner of the present application, the specific calculation process is as Figure 4 shown. Compared with the calculation method in the prior art, in the solution of the present application, the two separate LoRA models can be spliced and only need to be calculated once. At the same time, after calculating fused_LoRA_A_out, it is necessary to multiply by the routing weight fused_weight.

[0071] Figure 5 FIG. is a flowchart of a LoRA weight fusion method for a large language model inference system according to another exemplary embodiment. Figure 5 The process 50 shown is a detailed description of the Figure 1 way of dimension splitting in the process shown.

[0072] As Figure 5 shown, in S502, when the rank dimension of the LoRA weight data is greater than the dimension threshold, the LoRA weight data is split, and split into multiple LoRA processing weights.

[0073] First, the rank dimension of the LoRA weight data is judged. If its rank dimension is greater than the set dimension threshold, the LoRA weight data is split, and split into multiple smaller LoRA processing weight matrices.

[0074] Through the splitting process, the complexity of a single weight matrix is reduced, the storage and computing resources required for inference calculation are reduced, and at the same time, a foundation for subsequent flexible management is laid.

[0075] In S504, after the splitting process, multiple LoRA processing weights and their corresponding fusion ratios are stored based on the paging management method to generate the LoRA fusion weights. After the splitting is completed, the paging management method is used to effectively store and manage multiple LoRA processing weights and their corresponding fusion ratios.

[0076] The advantage of paging management can achieve dynamic loading of weights and allocate computing resources on demand. In a multi-task scenario, it supports quick switching between different weight combinations.

[0077] In S506, the input data is respectively subjected to inference calculations with multiple LoRA processing weights and their corresponding fusion ratios to generate LoRA results. The system respectively performs inference calculations on the input data with multiple LoRA processing weights and their corresponding fusion ratios to gradually generate their respective LoRA calculation results.

[0078] In a specific embodiment, the LoRA model to be fused can be split into several LoRA models of equal rank size in the rank dimension and combined with paging management for calculation.

[0079] Based on the paging management method, a LoRA model with a larger rank can be divided into several LoRA models with smaller ranks and stored in different pages respectively.

[0080] In the embodiments of the present application, according to paging storage Figure 6 As shown, they are respectively stored in three pages. Since the weights stored in each page must belong to the same LoRA model, when storing LoRA_weight, the shape can be defined as (num_pages, 1).

[0081] The specific calculation process is as Figure 7 shown. The results corresponding to the LoRA model in each page are calculated respectively, and then a final summation is performed. It is worth mentioning that although the following calculation process is calculated separately, it can be calculated in parallel during GPU calculation.

[0082] In the present application, although the two methods of splitting and splicing seem to be opposite, in fact, both can accelerate the calculation. This is because when fusing and calculating multiple LoRA models, the ranks of the LoRA models may vary, and whether multiple LoRAs need to be fused during the inference of the LLM system is dynamically determined (some inference requests require fusing multiple LoRA models, while some do not). In the technical solution provided by the present application, the splitting or splicing method can be flexibly selected according to the actual scenario for dynamic fusion.

[0083] In this application, the core similarity of the two methods lies in unifying the calculation process: Based on the splicing method, when multiple LoRA models need to be fused, they are spliced into a large LoRA model in the rank dimension. By uniformly processing all LoRA models (such as padding to the same rank), the unification of the calculation process is achieved.

[0084] Based on the segmentation method, whether fusion is required or not, each LoRA model is segmented into sub-LoRA models of equal size in the rank dimension, thus achieving the unification of calculation.

[0085] Although the two methods are different in operation methods, their common goal is to optimize the fusion calculation of multiple LoR models through a unified calculation process.

[0086] Those skilled in the art can understand that all or part of the steps to implement the above embodiments are realized as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above methods provided in this application are executed. The program can be stored in a computer-readable storage medium, which can be a read-only memory, a disk, an optical disc, etc.

[0087] In addition, it should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the methods according to the exemplary embodiments of this application, rather than for restrictive purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0088] The following is an embodiment of the device of this application, which can be used to execute the method embodiment of this application. For details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.

[0089] Figure 8 It is a block diagram of a LoRA weight fusion device for a large language model inference system shown according to an exemplary embodiment. As Figure 8 shown, the LoRA weight fusion device 80 for a large language model inference system includes: a data module 802, a proportion module 804, a processing module 806, an input module 808, and an inference module 810.

[0090] The data module 802 is used to obtain multiple LoRA weight data for the large language model inference system; The proportion module 804 is used to determine the fusion proportion corresponding to the multiple LoRA weight data; the proportion module 804 is also used to determine the fusion proportion corresponding to the multiple LoRA weight data according to the preset task requirements.

[0091] The processing module 806 is used to splice or split the multiple LoRA weight data based on the fusion ratio to generate LoRA fused weights; the processing module 806 is further used to splice or split the multiple LoRA weight data based on the rank dimension of the LoRA weight data; and fuse the data after splicing or splitting according to the fusion ratio to generate the LoRA fused weights.

[0092] The input module 808 is used to obtain input data for the inference system of the large language model; The inference module 810 is used to perform inference calculations by invoking the LoRA fused weights based on the input data. The inference module 810 is further used to perform inference calculations on the input data and the fused weights to generate LoRA results; and add the LoRA results and the output results of the large language model to generate inference results.

[0093] The LoRA weight fusion device for the large language model inference system according to the present application obtains multiple LoRA weight data through the large language model inference system; determines the fusion ratio corresponding to the multiple LoRA weight data; splices or splits the multiple LoRA weight data based on the fusion ratio to generate LoRA fused weights; the inference system of the large language model obtains input data; and performs inference calculations by invoking the LoRA fused weights based on the input data. This method is applicable to scenarios that require dynamic multi-style fusion, can improve the video memory utilization rate, and meet the dual requirements of flexibility and performance of the large language model inference system.

[0094] Figure 9 It is a block diagram of an electronic device shown according to an exemplary embodiment.

[0095] Next, refer to Figure 9 to describe the electronic device 900 according to this embodiment of the present application. Figure 9 The shown electronic device 900 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0096] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include but are not limited to: at least one processing unit 910, at least one storage unit 920, a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910), a display unit 940, etc.

[0097] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 910 can execute as Figure 1 , Figure 2 , Figure 5 shown in.

[0098] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202, and may further include a read-only storage unit (ROM) 9203.

[0099] The storage unit 920 may further include a program / utilities 9204 having a set (at least one) of program modules 9205. Such program modules 9205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0100] The bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0101] The electronic device 900 may also communicate with one or more external devices 900' (such as a keyboard, a pointing device, a Bluetooth device, etc.), so that the device that enables a user to interact with the electronic device 900 communicates, and / or the electronic device 900 can communicate with any device that can communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 950. And, the electronic device 900 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 960. The network adapter 960 can communicate with other modules of the electronic device 900 through the bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0102] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a manner of software combined with necessary hardware. Therefore, as Figure 10As shown, the technical solution according to the embodiment of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a portable hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present application.

[0103] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0104] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0105] The program code for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0106] The above computer-readable medium carries one or more programs, which, when executed by the device, cause the computer-readable medium to implement the following functions: the large language model inference system obtains a plurality of LoRA weight data; determines the fusion ratios corresponding to the plurality of LoRA weight data; splices or slices the plurality of LoRA weight data based on the fusion ratios to generate LoRA fused weights; the inference system of the large language model obtains input data; and performs inference calculations by invoking the LoRA fused weights based on the input data.

[0107] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices different from the embodiments. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.

[0108] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0109] The above specifically shows and describes the exemplary embodiments of the present application. It should be understood that the present application is not limited to the detailed structures, settings, or implementation methods described here; on the contrary, the present application is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. A LoRA weight fusion method for a large language model inference system, characterized in that, including: The large language model inference system obtains multiple LoRA weight data; Determine the fusion ratios corresponding to the multiple LoRA weight data; Based on the fusion ratios, splice or split the multiple LoRA weight data to generate LoRA fusion weights; The inference system of the large language model obtains input data; Based on the input data, call the LoRA fusion weights for inference calculation.

2. The method according to claim 1, characterized in that, Determine the fusion ratios corresponding to the multiple LoRA weight data, including: Determine the fusion ratios corresponding to the multiple LoRA weight data according to preset task requirements.

3. The method according to claim 1, characterized in that Based on the fusion ratios, splice or split the multiple LoRA weight data to generate LoRA fusion weights, including: Splice or split the multiple LoRA weight data based on the rank dimension of the LoRA weight data; Fuse the data after splicing or splitting according to the fusion ratios to generate the LoRA fusion weights.

4. The method according to claim 3, wherein Splice or split the multiple LoRA weight data based on the rank dimension of the LoRA weight data to generate the LoRA processing weights, including: When the rank dimension of the LoRA weight data is less than the dimension threshold, splice the multiple LoRA weight data and summarize them into the LoRA processing weights.

5. The method according to claim 4, characterized in that, Fuse the data after splicing or splitting according to the fusion ratios to generate the LoRA fusion weights, including: According to the fusion ratios, perform matrix splicing on the LoRA processing weights according to the rank dimension to generate the LoRA fusion weights.

6. The method according to claim 3, wherein Splice or split the multiple LoRA weight data based on the rank dimension of the LoRA weight data to generate the LoRA processing weights, including: When the rank dimension of the LoRA weight data is greater than the dimension threshold, split the LoRA weight data into multiple LoRA processing weights.

7. The method according to claim 6, wherein Fuse the data after splicing or splitting according to the fusion ratios to generate the LoRA fusion weights, further including: After performing the splitting process, store multiple LoRA processing weights and their corresponding fusion ratios based on the paging management method to generate the LoRA fusion weights.

8. The method according to claim 1, wherein Based on the input data, call the LoRA fusion weights for inference calculation, including: Perform inference calculation on the input data and the fusion weights to generate LoRA results; Add the LoRA results and the output results of the large language model to generate inference results.

9. The method according to claim 8, wherein Perform inference calculation on the input data and the fusion weights to generate LoRA results, including: After performing the splitting process, perform inference calculation on the input data respectively with multiple LoRA processing weights and their corresponding fusion ratios to generate LoRA results.

10. A LoRA weight fusion device for a large language model inference system, characterized in that, including: A data module for the large language model inference system to obtain multiple LoRA weight data; A ratio module for determining the fusion ratios corresponding to the multiple LoRA weight data; A processing module for splicing or splitting the multiple LoRA weight data based on the fusion ratio to generate LoRA fusion weights; An input module for obtaining input data for the inference system of the large language model; An inference module for performing inference calculations by calling the LoRA fusion weights based on the input data.

Citation Information

Patent Citations

  • Large anesthesia model training method and device

    CN117095827A

  • Method for enhancing memory ability of large model to external knowledge base

    CN117114012A

  • Depth fine tuning method for generating Chinese text logical reasoning thinking chain

    CN117669536A

  • Large language model fine tuning and Adapter fusion method and device

    CN117708307A

  • Pre-training language model parameter fine tuning method and device, equipment and medium

    CN117829240A