Large Language Model Optimization System and Method Based on Low-Rank Adaptation Technology
Through a large language model optimization system based on low-rank adaptive technology, the problems of low inference efficiency and high resource consumption of large language models in the existing technology are solved, and the data volume and retrieval efficiency are reduced, and operation instructions are processed more accurately.
Patent Information
- Application Number
- CN202510337171.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-21
AI Technical Summary
When handling complex tasks, existing large language models have low inference efficiency and high resource consumption, especially in retrieval data generation and task recognition, and cannot effectively optimize instruction accuracy.
A large language model optimization system based on low-rank adaptive technology is adopted. The system includes a low-rank adjustment module, an instruction preliminary analysis module, a data comparison module, an instruction optimization module and a model evaluation module. Through low-rank decomposition and instruction optimization techniques, the calculation complexity and memory usage of the model are reduced, and the accuracy and efficiency of instructions are improved.
It effectively reduces the amount of data required to be processed, improves the efficiency of data retrieval, determines the purpose of operation instructions more accurately, reduces the internal consumption of the system, and achieves the effect of fundamentally solving the problem of excessive data.
Smart Images

Figure CN119886279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and specifically to an optimization system and method for large language models based on low-rank adaptation technology. Background Art
[0002] In recent years, thanks to the support of big data and the enhancement of computing power, artificial intelligence technology has developed rapidly, and large language models based on generative artificial intelligence technology have emerged. It performs excellently in simulating human thinking logic and language processing capabilities, laying a foundation for the development of intelligent assistance tools and intelligent assistants. All walks of life have also begun to explore the application of large language models (LLMs) in their respective industries. However, with the wide application of large language models, inference efficiency and resource consumption have become the main bottlenecks restricting their expansion in practical applications, especially in complex tasks such as retrieving data generation and task recognition.
[0003] Currently, a Chinese patent with the patent application number "CN202410402548.2" discloses a fine-tuning method for large language models based on low-rank matrix decomposition, including obtaining the pre-trained weight file of the large language model and the question-and-answer pairs required for fine-tuning the large language model; using a double quantization method to compress the precision of the model parameters in the pre-trained weight file; replacing the fully connected layer in the large language model structure with a LoRA layer; performing two rounds of dequantization on the precision-compressed model parameters in batches and calculating the output of the LoRA layer; in each batch, calculating the loss based on the output of the LoRA layer and dynamically updating the parameters of the LoRA layer based on the loss backpropagation until all batches are processed, and outputting the fine-tuned large language model. Although it can master knowledge in multiple fields by setting multiple parallel LoRA modules, and then use a routing module to direct the input question to the LoRA module responsible for that field to obtain the output, mainly in terms of how to improve the data processing speed, for example, as described above, by dividing the data required for the operation into multiple independent branch operation data and then performing synchronous operations by multiple computing devices. However, it ignores how to optimize the accuracy of the obtained instructions.
[0004] In addition, the Chinese patent with the existing patent application number "CN202411045654.6" discloses a large language model acceleration method and device, which has the following steps: S1, compress and decompose the pre-trained weight matrix W of the large language model; S2, perform QR decomposition on the U and VT matrices respectively to obtain QU, RU, QV, RV; S3, use the singular value decomposition algorithm to compress and decompose the matrix product; S4, merge the matrices respectively; S5, replace the weight matrix W with the obtained matrix for storage; S6, use the stored matrix for reasoning acceleration; S7, set the obtained rank r to the rank of the low-rank parameterized update matrix ΔW corresponding to the weight matrix W; S8, use the stored matrix to fine-tune the reasoning acceleration of the large language model. It also focuses on how to improve the speed of the operation itself, but whether it is through improving the operation speed of the equipment or adjusting the operation method, it will take a lot of time and energy for research and development, and it is necessary to find another way.
[0005] However, during the implementation of the above technical solution, it was found that there were at least the following technical problems:
[0006] Unable to optimize the instructions to be executed: The existing large language model improvements are mainly divided into two categories (hardware and software). One is to improve the computing power of the computing device by hardware structure, for example, the processor Intel Core i9-14900K and the processor Intel Core i7-14700K, the i9 has 16 cores and 32 threads, while the i7 has 8 cores and 16 threads. It can be seen that the i9 can handle multiple tasks at the same time compared to the i7, reducing the CPU occupancy rate to improve computing power. This method requires a lot of time and energy to invest in research and development, and the research and development is difficult. At the same time, due to the inability to filter or optimize the data, the amount of data that needs to be processed remains unchanged, and this problem cannot be fundamentally solved; the other type is to reduce the amount of data that needs to be calculated by splitting and adjusting the retrieval instructions. However, no matter what calculation method is used, it is necessary to process and execute according to the received retrieval instructions. For example, when issuing a historical data retrieval instruction, since it is impossible to accurately determine which historical data is required, all data can only be retrieved. After system optimization, only the previous retrieval habits are added, so that the data is divided into before and after or important or not. However, as time goes by, the data continues to accumulate, resulting in the data that needs to be processed is still very large, and the data optimization has reached a bottleneck, so it is difficult to have a new breakthrough. To sum up, it can be seen that both hardware and software improvements are very difficult and have limited effects. Therefore, we need to open up a new way to optimize and adjust large language models. Summary of the invention
[0007] (1) Technical problems to be solved
[0008] In view of the deficiencies of the prior art, the present invention provides a large language model optimization system and method based on low-rank adaptation technology, which solves the problems raised in the background art.
[0009] (2) Technical solutions
[0010] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0011] A large language model optimization system based on low-rank adaptation technology, the optimization system includes:
[0012] A low-rank adjustment module, which uses the input data that has been cleaned and preprocessed as the dataset for training the large language model, trains the large language model, and then performs low-rank decomposition and optimization on the initial model by low-rank adaptation technology. Among them, low-rank decomposition is to decompose the original high-dimensional parameter matrix into the product of two or more low-rank matrices;
[0013] An instruction preliminary analysis module, which extracts the keywords in the received initial instruction as retrieval terms, obtains the retrieval data that can be called by the large language model, and divides them into corresponding retrieval datasets according to types. Then, according to the number of subsets in the retrieval dataset, it is divided into a preliminary retrieval dataset and a discardable retrieval dataset, and the discardable retrieval dataset is discarded;
[0014] A data comparison module, when receiving the preliminary retrieval dataset, sorts the preliminary retrieval datasets in descending order according to the number of subsets in the preliminary retrieval dataset, selects the first several preliminary retrieval datasets as the preferred datasets, and compares the proportion of the preferred datasets in all preliminary retrieval datasets with the preset screening ratio threshold. When the proportion is less than the screening ratio threshold, the preliminary retrieval datasets are output as the retrieval results in the arranged order; otherwise, the comparison data is output as the first retrieval result, and the estimated impact index of the instruction execution is analyzed and checked, and a check instruction is issued;
[0015] An instruction optimization module, when receiving the check instruction, performs secondary classification on the preferred datasets according to the classification criteria stored in the database, analyzes the data classification evaluation value of the secondary classification, combines it with the historical retrieval trend to select the corresponding classification criteria as the optimized instruction output, and then records the instruction after the optimized instruction feedback as the secondary instruction, and screens and outputs the retrieved data after the secondary classification;
[0016] A model evaluation module, which obtains the proportion of the preferred datasets in the preliminary retrieval datasets during the retrieval process, the estimated impact index, and the number of subsets of the preferred datasets before and after the instruction optimization, generates a model optimization efficiency evaluation coefficient, and when the model optimization efficiency evaluation coefficient is lower than the preset optimization effect threshold, retrieves the optimization scheme correction strategy.
[0017] Furthermore, the input data passes through all training layers, and the output of the training layer is combined with the input adjusted by the new parameters to generate a new adjusted output; and in the case of low-rank decomposition, importance-based pruning or structure-based pruning is used to remove unimportant parameters in the model.
[0018] Furthermore, the subset of the search data set represents the search data that can be directly obtained according to the search terms. When the search data set is classified, it is compared with the preset rejection threshold. When the number of subsets in the search data set is less than the preset rejection threshold, it is recorded as a discardable search data set; otherwise, it is recorded as a preliminary search data set.
[0019] Among them, when searching according to the secondary instruction, if the number of subsets in the discardable search data set increases, and the ratio of the increase to the number of atomic sets is greater than 20%, a reuse instruction is issued to record the corresponding discardable search data set as a preliminary search data set; otherwise, the corresponding discardable search data set is recorded as a non-search data set under the search term, and is not displayed in the callable search data retrieved by the same search term.
[0020] Furthermore, the process of adjusting the discard threshold is as follows:
[0021] Set the adjustment ratio, denoted as ;
[0022] Get the number of discardable search datasets under the current search term and the number of preliminary search datasets, calculate the ratio of the discardable search dataset to the number of preliminary search datasets, denoted as , and The discard ratio threshold is compared with the preset discard ratio threshold, wherein the discard ratio threshold includes the discard upper threshold and the discard lower threshold; when >Discard the upper threshold or When the lower threshold is discarded, an adjustment instruction is issued and the adjustment strategy is executed; otherwise, no response is made;
[0023] When receiving the adjustment instruction, the adjustment strategy is executed to calculate the average number of subsets of all retrieved data sets, which is recorded as , when the average number of subsets >Discard upper threshold× , an upward command is issued to × As the adjusted discard upper threshold;
[0024] When the average number of subsets <Discard lower threshold× , a downward adjustment command is issued to Discard the lower threshold as the adjusted lower discard threshold; otherwise, do not respond.
[0025] Further, the process of analyzing the proportion of the preferred dataset is as follows:
[0026] Retrieve the average viewing volume, the current viewing volume, and the current total viewing duration under the previous n retrieval instructions from the database;
[0027] Use the average viewing volume as the quantity selected from the preliminary retrieval dataset. The selected preliminary retrieval dataset is denoted as the preferred dataset. Among them, when the total number of the preliminary retrieval dataset ≤ the average viewing volume, an early output instruction is issued, and all the preliminary retrieval datasets are output as the retrieval results; otherwise, compare the proportion of the preferred dataset in all the preliminary retrieval datasets with the preset screening ratio threshold.
[0028] When receiving the early output instruction, if the current viewing volume ≥ the average viewing volume and the current total viewing duration ≤ the preset viewing duration, a call instruction is issued to retrieve the discardable retrieval dataset as the non-conventional retrieval dataset for output; otherwise, generate a residence time distribution map of each retrieval dataset at the current viewing volume, and use the retrieval datasets that do not reach the preset standard duration as the screening criteria, and remove the retrieval datasets with the same classification as the screening criteria from the unviewed retrieval datasets.
[0029] Further, the process of analyzing the estimated influence index is as follows:
[0030] Retrieve the number of discardable retrieval datasets and the number of discardable retrieval datasets converted into preliminary retrieval datasets during the previous L retrievals, denoted as 、 ;
[0031] Through the analysis formula Obtain the average conversion ratio , where represents the number of discardable retrieval datasets at the m-th retrieval, represents the number of discardable retrieval datasets converted into preliminary retrieval datasets at the m-th retrieval;
[0032] Schedule the number of preliminary retrieval datasets, discardable retrieval datasets, and their subsets, denoted as 、 、 、 ;
[0033] Through the analysis formula Obtain the estimated influence index , where represents the time required for data verification, Represents the total number of preliminary retrieval datasets, Represents the total number of discardable retrieval datasets, Represents the i-th preliminary retrieval dataset, Represents the number of subsets within the i-th preliminary retrieval dataset, Represents the j-th discardable retrieval dataset, Represents the number of subsets within the j-th discardable retrieval dataset, Represents the speed of identifying the number of unit data sets, Represents the correction factor of the preset estimated impact index, Represents the total data volume of the dataset, Represents the actual time required for the previous execution of the same verification instruction, Represents the estimated impact index for the previous execution of the same verification instruction, Represents the time difference between the actual time required for the previous execution of the same verification instruction and the time required for data verification, Respectively represent the weights of the preset time required for data verification, the total data volume of the dataset, and the actual time required for the previous verification instruction, and , where e represents the natural constant;
[0034] When the estimated impact index > the preset impact risk threshold, a data output instruction is issued, and the preliminary retrieval datasets are output as retrieval results in the arranged order; otherwise, a qualified instruction is issued to start executing the verification instruction.
[0035] Furthermore, the historical retrieval trend represents the number of selections of various classification criteria under the same retrieval terms. The analysis process of the data classification evaluation value of the secondary classification is as follows:
[0036] Retrieve the classification criteria stored in the database to perform secondary classification on the preferred dataset, and obtain the classification sets after classification. Among them, the preferred dataset is a subset of the classification set, and obtain the two classification sets with the largest and smallest number of subsets. The number of subsets in the two classification sets is respectively recorded as 、 ;
[0037] When ≥ the preset uniformity threshold, a classification unqualified instruction is issued, and secondary classification is performed again; otherwise, a classification qualified instruction is issued, and the classification criteria corresponding to this secondary classification are recorded;
[0038] When receiving the classification qualified instruction, obtain all the classification sets and the corresponding number of subsets after secondary classification, and record them as 、 , and retrieve the two classification sets with the largest and smallest number of subsets based on this classification criteria, where the number of subsets is respectively recorded as , ;
[0039] By analyzing the formula:
[0040] Obtain the data classification evaluation value of the secondary classification , where represents the total number of classification sets after secondary classification, represents the m-th classification set, represents the number of subsets corresponding to the m-th classification set, represents the number of subsets in the m-th classification set, respectively represent the weights of the preset time required for secondary classification and the degree of uniformity of classification sets, ;
[0041] By analyzing the formula Obtain the scoring coefficient of the classification criterion , where represents the number of selections of the y-th classification criterion, represents the data classification evaluation value of the y-th classification criterion, represents the weights of the preset number of selections and the data classification evaluation value, , and according to the scoring coefficient of the classification criterion Select the corresponding classification criteria in descending order as the optimization instruction to be issued.
[0042] Furthermore, the formula on which the model optimization efficiency evaluation coefficient is based is as follows:
[0043] where represents the model optimization efficiency evaluation coefficient, Y represents the total number of preferred data sets before the optimization instruction, represents the k-th preferred data set before the optimization instruction, represents the number of subsets corresponding to the k-th preferred data set before the optimization instruction, X represents the total number of preferred data sets after the optimization instruction, represents the x-th preferred data set after the optimization instruction, represents the number of subsets corresponding to the x-th preferred data set before the optimization instruction, respectively represent the weights of the preset proportion of preferred data sets in the preliminary retrieval data set, the estimated influence index, and the ratio of the number of subsets of preferred data sets before and after instruction optimization, , and when the total number of preliminary retrieval data sets ≤ average viewing volume, It is recorded as 1.
[0044] Further, the correction strategy is activated when the model optimization efficiency evaluation coefficient is lower than the optimization effect threshold. The corresponding classification criteria are sequentially selected in the order of the scoring coefficient from large to small, and the corresponding model optimization efficiency evaluation coefficients are calculated respectively. Then, they are compared with the optimization effect threshold, and the classification criteria corresponding to the model optimization efficiency evaluation coefficient lower than the preset optimization effect threshold are selected as the optimization scheme output.
[0045] Further, a large language model optimization method based on low-rank adaptation technology, the optimization method includes the following steps:
[0046] Taking the input data that has been cleaned and preprocessed as the training dataset of the large language model, training the large language model, and then performing low-rank decomposition and optimization on the initial model by low-rank adaptation technology, where low-rank decomposition is to decompose the original high-dimensional parameter matrix into the product of two or more low-rank matrices;
[0047] Extracting the keywords in the received initial instruction as retrieval entries, obtaining the retrieval data that can be called by the large language model, and classifying them into corresponding retrieval datasets according to the type. Then, they are divided into a preliminary retrieval dataset and a discardable retrieval dataset according to the number of subsets in the retrieval dataset, and the discardable retrieval dataset is discarded;
[0048] When receiving the preliminary retrieval dataset, sorting according to the number of subsets in the preliminary retrieval dataset from large to small, selecting the first several preliminary retrieval datasets as the preferred dataset, and comparing the proportion of the preferred dataset in all preliminary retrieval datasets with the preset screening proportion threshold. When the proportion is less than the screening proportion threshold, the preliminary retrieval dataset is output as the retrieval result in the arranged order; otherwise, the comparison data is output as the first retrieval result, and the estimated impact index of the instruction execution is analyzed and a verification instruction is issued;
[0049] When receiving the verification instruction, performing secondary classification on the preferred dataset according to the classification criteria stored in the database, analyzing the data classification evaluation value of the secondary classification, combining it with the historical retrieval trend to select the corresponding classification criteria as the optimization instruction output, and then recording the instruction after the optimization instruction feedback as the secondary instruction, and screening and outputting the retrieved data after the secondary classification;
[0050] Obtaining the proportion of the preferred dataset in the preliminary retrieval dataset, the estimated impact index, and the number of subsets of the preferred dataset before and after instruction optimization during the retrieval process, generating a model optimization efficiency evaluation coefficient, and when the model optimization efficiency evaluation coefficient is lower than the preset optimization effect threshold, invoking the optimization scheme correction strategy.
[0051] (III) Beneficial effects
[0052] The present invention provides a large language model optimization system and method based on low-rank adaptation technology, having the following beneficial effects:
[0053] Using the large language model optimized by low-rank adaptation technology as the basis for data processing, while facilitating subsequent optimization, it can effectively reduce the base of the data volume to be processed; when executing operation instructions, according to the proportion of the number of retrieved contents in the directly obtained retrieval dataset, part of the data is discarded, thereby reducing the base volume of subsequent retrieved data, and thus improving the efficiency of data retrieval; in addition, according to the proportion of the preferred dataset in all preliminary retrieval datasets, a "question" is sent to the "input end" to optimize the operation instructions, avoiding the "internal consumption" of the system, and through this "questioning" method, the purpose of the operation instructions can be determined more accurately, so as to give an accurate feedback, and the total amount of data to be processed is reduced to a greater extent, so as to fundamentally solve the problem of excessive data to be processed.
[0054] Before sending the "question", the following two steps are performed: First, analyze the verification instructions to be performed to obtain an estimated influence index that can predict the difficulty of executing the verification instructions, so as to conveniently avoid those cumbersome and complex verification instructions in advance and improve the efficiency of data analysis; Second, use the classification criteria stored in the database as the secondary classification criteria for the preferred dataset to divide the preferred dataset into different categories, and then calculate and generate a classification evaluation value that can evaluate whether the classification criteria are appropriate and the degree of appropriateness based on the obtained classification results, so as to objectively evaluate the classification criteria, and combine it with the historical retrieval trend to select the classification criteria that meet the requirements as the optimized instruction output, providing options for the feedback of the instructions, so as to reduce the feedback time.
[0055] Using the proportion of the preferred dataset in the preliminary retrieval dataset, the estimated influence index, and the number of subsets of the preferred dataset before and after instruction optimization during the process of sending a "question" to the "input end", a model optimization efficiency evaluation coefficient for intuitively evaluating the improvement of the efficiency of this retrieval optimization method is obtained, so as to have a clear understanding of the entire optimization plan, providing a basis for orderly improvement and adjustment, and then through the correction of the correction strategy, an optimization plan that meets the optimization requirements is obtained. Brief Description of the Drawings
[0056] Figure 1 It is the overall structure flowchart of the present invention;
[0057] Figure 2 It is the schematic diagram of the model evaluation process of the present invention;
[0058] Figure 3 It is the distribution diagram of the residence time of each retrieval dataset under the current view volume in the present invention. Detailed Embodiment
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Low-Rank Adaptation (LoRA) is a technique for fine-tuning large language models (LLMs), aiming to solve the problem of huge computational resources and time consumption when fine-tuning such models. The LoRA technique significantly reduces the number of parameters to be fine-tuned by freezing the weights of the pre-trained model and injecting trainable rank decomposition matrices in each Transformer block, thereby reducing the computational cost and improving the training efficiency.
[0061] With the wide application of large language models (LLMs), inference efficiency and resource consumption have become the main bottlenecks restricting their expansion in practical applications. Especially in complex tasks such as text-to-image prompt generation, optimizing memory usage and inference speed becomes particularly important. However, through research, it is found that existing large language models spend a lot of time during inference and analysis, and the results obtained are very different from what users need. For the sake of easy understanding, the time and proportion obtained during the inference ability test are represented as integers (rounded).
[0062] Inference ability test: Randomly send retrieval instructions to the large language model, and the interval between two adjacent instructions needs to be controlled within 10 - 15 s. Then record the time required for the response and the total time required after the retrieval is completed, and record the proportion of the content required in the result feedback by the large language model. The results show that the response time is generally between 3 - 8 s, and there are obvious differences in the same content due to the different number of keywords or different content in the retrieval instructions. For the sake of easy understanding, the following examples are all based on conventional text retrieval. For example, it is the 18th in Guangzhou on that day. If the query instruction is "Get the historical temperature in Guangzhou", the result obtained is the temperature change values from the 1st to the 17th in Guangzhou, and the response time is 3 s (including the response and recognition time of the device); if the query instruction is "Get the temperature from the 13th to the 17th in Guangzhou", the response time is only 1 s.
[0063] After analysis, it can be seen that this is because when the large language model conducts reasoning and analysis, due to inaccurate received instructions, the time required for independent reasoning varies. During the independent reasoning process, in order to ensure that the information is what the user needs, it is necessary to comprehensively obtain it by combining the user's retrieval situation, coordinates, etc. This increases the data base for reasoning, seriously affecting its response efficiency, and there is a large amount of unnecessary content in the obtained data. The existence of this unnecessary content is also because the model cannot analyze and reason whether this content is what the user needs, which also reflects from the side the inaccuracy of the input instructions, resulting in complex obtained data and easily affecting the user's search.
[0064] However, the existing optimization methods still stay at how to improve the data processing speed. For example, by dividing the data required for operations into multiple independent branch operation data and then performing synchronous operations by multiple operation devices. But it ignores how to optimize the accuracy of the obtained instructions.
[0065] Summary: Using the large language model optimized by low-rank adaptation technology as the basis for data processing, and then according to the proportion of the preferred data set in all preliminary retrieval data sets, send a "question" to the "input end" to optimize the operation instructions, avoid the "internal consumption" of the system, and through this "questioning" method, the purpose of the operation instructions can be determined more accurately, so as to give accurate feedback, and minimize the total amount of data to be processed to fundamentally solve the problem of excessive data to be processed.
[0066] Research and development concept:
[0067] Since the initial instructions for the large language model to reason and analyze are issued by the user and have practicality and uncontrollability, it is necessary to obtain the accurate needs of the user. For this reason, we use the form of questions and answers to refine the instructions, and analyze the preliminary instructions during the process of the document. For example, when retrieving the "working principle of air conditioners", the data obtained by the large language model is mainly divided into two types: "refrigeration principle of air conditioners" and "heating principle of air conditioners". Before content generation, a verification instruction can be issued, and it is issued in the form of a question, such as, do you need the "refrigeration principle of air conditioners" or the "heating principle of air conditioners", and use the obtained feedback instruction as the secondary instruction for execution. In this way, no matter whether the "refrigeration principle of air conditioners" or the "heating principle of air conditioners" is selected, the first issued instruction can be optimized. In addition, the two materials can be prepared simultaneously during the interval of the questions to facilitate data extraction.
[0068] At the same time, to improve efficiency, when the ratio of retrieving the "refrigeration principle of air conditioners" and the "heating principle of air conditioners" exceeds the set threshold, the result with the larger proportion is used as the output result; otherwise, verification is carried out.
[0069] Example 1:
[0070] Basic data optimization
[0071] Clean and preprocess the data, and use it as the dataset for training the big data model. Train the large language model, and then use the low-rank adaptation technology to perform low-rank decomposition and optimization on the initial model. Among them, low-rank decomposition is to decompose the original high-dimensional parameter matrix into the product of two or more low-rank matrices, that is, to decompose a large matrix into the product of two or more smaller and simpler matrices, and these small matrices usually have lower ranks. This decomposition method helps to reduce the number of parameters, lower the computational complexity and memory occupancy, and is often used for model compression and acceleration.
[0072] Among them, in the traditional Adapter, by inserting relatively small trainable parameters between layers, the number of trainable parameters of the model is reduced. However, this strategy has a significant defect, that is, there is no direct path for gradients during backpropagation, and the gradients need to pass through all training layers. Therefore, this structure also faces certain challenges in parallel computing. Similarly, Prompt Tuning also faces a similar problem, that is, the lack of a direct gradient path, resulting in the gradients passing through all training layers during backpropagation. In addition, the task setting of Prompt Tuning is relatively idealized, trying to optimize the model only by adjusting a small number of parameters at the input end. This method has limited impact on the deep part of the model, thus limiting the effect of fine-tuning.
[0073] To address the above problems, an improved method of low-rank coding optimizes the Transformer structure. This method first passes the input data through all training layers, then combines the output of the training layers with the input adjusted by the new parameters, and finally generates a new adjusted output. In this way, the Transformer structure optimized by low-rank coding only needs to optimize the new parameters without loading and training the original parameters. The improved low-rank coding performs transfer learning on the ladder parameter memory based on LoRA (low-rank adaptation), and can effectively learn and fuse the reduced rank and the target data more fully.
[0074] Execute the preliminary instructions
[0075] "Initial instruction" refers to the first issued instruction. After receiving the initial instruction, extract the keywords therein as retrieval terms, and retrieve them in the retrieval data that can be called by the large language model. Then, classify them into corresponding retrieval data sets according to the type. This classification type is the secondary classification type retrieved under the same retrieval terms last time (the secondary classification will be introduced later). Then, divide them into the initial retrieval data set and the discardable retrieval data set according to the number of subsets in the retrieval data set, and discard the discardable retrieval data set. Among them, the subset of the retrieval data set refers to the retrieval data that can be directly obtained according to the retrieval terms.
[0076] When classifying the retrieval data set, it is obtained by comparing the number of subsets in the retrieval data set with a preset discard threshold. When the number of subsets in the retrieval data set is less than the preset discard threshold, it is recorded as the discardable retrieval data set; otherwise, it is recorded as the initial retrieval data set. The reason for adopting this method is that for some data with very small quantities, the content recorded in them is limited and not universal. Suppose the content to be retrieved is: "Power station address", and there are 10 pieces of data related to power station one, 15 pieces of data related to power station two, and 2 pieces of data related to power station three. It can be seen that for power station three, the relevant information is very little (indicating that it has little attention or is less studied as an object), and it is not concerned. Therefore, this kind of data can be discarded to reduce the retrieval volume of the overall data.
[0077] Although these data have been discarded as unimportant data, in order to avoid special needs, it is still necessary to recycle these data. For example, when retrieving according to the secondary instruction, if the number of subsets in the discardable retrieval data set increases, it indicates that the secondary instruction is more biased towards the part discarded in the first classification. When the ratio of the increased quantity to the original subset quantity > 20%, a reuse instruction is issued, and the corresponding discardable retrieval data set is recorded as the initial retrieval data set, and the content discarded in the first classification is called again to meet the needs of unconventional data; on the contrary, when the ratio of the increased quantity to the original subset quantity ≤ 20%, the corresponding discardable retrieval data set is recorded as the non-retrieval data set under this retrieval term and is not displayed in the retrievable retrieval data retrieved under the same retrieval term, reducing the basic data volume for calculation.
[0078] However, it is found in use that due to the different amounts of data obtained by different retrieval instructions, if a fixed-size discard threshold is used as the screening basis, it is very easy to have errors. Therefore, we propose an adjustment scheme for the discard threshold to be applied to changing data. The steps for adjusting the discard threshold are as follows:
[0079] The first step: Set a standard adjustment ratio , to limit the adjustment range of the discard threshold;
[0080] Step 2: Obtain the preliminary instruction, and calculate the number of discardable retrieval data sets based on the retrieval terms within the instruction, and the number of preliminary retrieval data sets. Calculate the ratio of the number of discardable retrieval data sets to the number of preliminary retrieval data sets, denoted as . Assume that the total number of retrievable data (i.e., all data retrievable with the preliminary instruction) is 100 data. After comparison with the discard threshold, the number of discardable retrieval data sets is 20, and the number of preliminary retrieval data sets is 100 - 20 = 80. Then the ratio of the number of discardable retrieval data sets to the number of preliminary retrieval data sets ; ;
[0081] Step 3: Compare with the preset discard ratio threshold range. The discard ratio threshold is divided into an upper discard threshold and a lower discard threshold. When > the upper discard threshold, it means that the proportion of discarded data in all retrieved data is very large. If all are discarded, the amount of basic data that can be provided will be very small. On the contrary, when < the lower discard threshold, it means that the proportion of discarded data is very small, and there is still a large amount of basic data to be analyzed. Therefore, whether > the upper discard threshold or < the lower discard threshold, adjustments need to be made. At this time, an adjustment instruction is issued and the adjustment strategy is executed. When the lower discard threshold ≤ ≤ the upper discard threshold, no adjustment is required;
[0082] Step 4: Adjustment strategy: Calculate the average subset number of all retrieval data sets, denoted as ; Upward adjustment: When the average subset number > the upper discard threshold × , an upward adjustment instruction is issued, and × the upper discard threshold is used as the adjusted upper discard threshold;
[0083] Downward adjustment: When the average subset number < the lower discard threshold × , a downward adjustment instruction is issued, and × the lower discard threshold is used as the adjusted lower discard threshold;
[0084] On the contrary, when the upper discard threshold × ≥ the average subset number ≥ the lower discard threshold × , no response is made.
[0085] Data volume comparison
[0086] Using the preliminary retrieval dataset obtained after preliminary instruction analysis as the basic data for analysis, first sort the obtained preliminary retrieval dataset in descending order according to the number of its internal subsets, and select the first several preliminary retrieval datasets from them as the preferred dataset. The basis for selecting the preferred dataset is the average viewing volume under the same retrieval instruction for the first n times. Suppose that only 10 pieces of data can be viewed per retrieval on average, then 10 pieces is the screening criterion. There are a total of 30 pieces of preliminary retrieval datasets obtained. Sort them according to the number of internal data, and select the first 10 pieces of preliminary retrieval datasets. These 10 pieces of preliminary retrieval datasets are the preferred dataset.
[0087] However, since the number of preliminary retrieval datasets is not fixed and may vary, when the total number of preliminary retrieval datasets ≤ the average viewing volume, an early output instruction is issued, and all preliminary retrieval datasets are output as the retrieval result; otherwise, the ratio of the preferred dataset in all preliminary retrieval datasets is compared with the preset screening ratio threshold. Among them, the comparison results are as follows:
[0088] When the ratio is less than the screening ratio threshold, it means that the preliminary retrieval datasets after preliminary instruction analysis are not enough to be selected, so the preliminary retrieval datasets can be directly output as the retrieval result according to the arranged order; on the contrary, when the ratio is greater than or equal to the screening ratio threshold, it means that the preliminary retrieval datasets after preliminary instruction analysis are very large and need to be refined. However, to avoid data generation delay, the comparison data is first output as the first retrieval result, so that there is no need to wait until the verification instruction is completed to obtain the result.
[0089] When receiving the early output instruction, if the current viewing volume ≥ the average viewing volume and the current total viewing duration ≤ the preset viewing duration, a call instruction is issued to retrieve the discardable retrieval dataset as the non-conventional retrieval dataset for output; otherwise, a residence time distribution map of each retrieval dataset under the current viewing volume is generated, and the retrieval datasets that do not reach the preset standard duration are used as the screening criterion to remove the retrieval datasets of the same classification as the screening criterion from the unviewed retrieval datasets. Suppose, Figure 3 For the distribution map, the Y-axis in the distribution map is time (residence time), and the X-axis is the number (the number of each retrieval dataset). Through Figure 3 It can be seen that the preview time mainly stays at number one and number four, and the time staying at the position of number three is less, probably between 15 - 20s. At this time, if the standard duration is 20s and the times of number three and number four in the figure are both less than 20s, then the data of the same type as number three and number four need to be used as the screening criterion to delete the remaining data that has not been previewed yet.
[0090] The above-mentioned average view count, current view count, and current total viewing duration are all obtained by retrieving data obtained under the same search instruction for the previous n times from the database. The average view count is obtained by adding up the view counts obtained in the previous n times and then dividing by n; the current view count represents the number of preferred data sets that have been viewed; the current total viewing duration represents the total time spent viewing the preferred data sets, and all these data are retrieved from the database.
[0091] When the verification instruction is issued, analyzing the verification instruction execution yields an estimated impact index that can predict the difficulty level of the verification instruction execution, thereby facilitating the avoidance of those cumbersome and complex verification instructions in advance and improving the efficiency of data analysis. The steps for conveniently selecting whether to execute the verification instruction based on the estimated impact index are as follows:
[0092] First, retrieve from the database the number of retrievable data sets that can be discarded during the previous L times of the same term search and the number of retrievable data sets that can be discarded and converted into a preliminary search data set .
[0093] (1.1)
[0094] In formula (1.1), represents the average conversion ratio of the retrievable data set that can be discarded and converted into a preliminary search data set, represents the number of retrievable data sets that can be discarded during the m-th search, represents the number of retrievable data sets that can be discarded and converted into a preliminary search data set during the m-th search.
[0095] (1.2)
[0096] (1.3)
[0097] In formulas (1.2) and (1.3), represents the estimated impact index for executing the verification instruction, represents the time required for data verification, represents the total number of preliminary search data sets, represents the total number of retrievable data sets that can be discarded, represents the number of preliminary search data sets, represents the number of subsets within the preliminary search data set, represents the i-th preliminary search data set, represents the number of subsets within the i-th preliminary search data set, represents the number of retrievable data sets that can be discarded, represents the number of subsets within the retrievable data set that can be discarded, Indicates the j-th discardable retrieval dataset, Indicates the number of subsets within the j-th discardable retrieval dataset, Indicates the speed of identifying the number of unit data sets, Indicates the correction factor for the preset estimated impact index, Indicates the total data volume of the dataset, Indicates the actual time required for the previous execution of the same verification instruction, Indicates the estimated impact index for the previous execution of the same verification instruction, Indicates the time difference between the actual time required for the previous execution of the same verification instruction and the time required for data verification, Respectively indicate the weights of the preset time required for data verification, the total data volume of the dataset, and the actual time required for the previous verification instruction, and , where e represents the natural constant.
[0098] When the estimated impact index > the preset impact risk threshold, it indicates that under the current classification standard, both the data to be operated on and the time required are very large. If executed, it will require a large amount of time and effort, seriously affecting the efficiency of system execution. Therefore, at this time, a data output instruction is issued, and the preliminary retrieval dataset is output as the retrieval result according to the arranged order, and the result is directly issued.
[0099] When the estimated impact index ≤ the preset impact risk threshold, a qualified instruction is issued, and the verification instruction is started to be executed. The current classification standard is combined with the historical retrieval trend to select the corresponding classification standard as the optimization instruction for output. Then, the instruction after the feedback of the optimization instruction is recorded as the secondary instruction, and the retrieved data after secondary classification is screened and output.
[0100] Instruction optimization
[0101] This is also the focus of this application. It mainly lies in "asking questions" to the "input end" according to the proportion of the preferred dataset in all preliminary retrieval datasets, so as to optimize the operation instructions, avoid the "internal consumption" of the system, and through this "questioning" method, the purpose of the operation instructions can be determined more accurately, so as to give accurate feedback, and the total amount of data to be processed can be reduced to a greater extent, so as to fundamentally solve the problem of too much data to be processed. Thus, some unclear instructions are refined to reduce the time for data analysis and data derivation.
[0102] Execution instruction optimization is enabled when a verification instruction is received. After enabling, according to the classification criteria stored in the database (such as the order of the first letter of the name, type, upload time, data volume, etc.), the preferred data set is classified again, and the data classification evaluation value of the secondary classification is analyzed. The data classification evaluation value is a classification evaluation value that can be calculated based on the classification results obtained after the secondary classification to evaluate whether the classification criteria are appropriate and the degree of appropriateness, so as to objectively evaluate the classification criteria. Combining it with the historical retrieval trend, the classification criteria that meet the requirements can be selected as the optimization instruction output, providing options for the feedback of the instruction, so as to reduce the feedback time. The formula on which the classification evaluation value is based is as follows:
[0103] (2.1)
[0104] In formula (2.1), represents the data classification evaluation value for secondary classification under the current classification criteria, represents the total number of classification sets after secondary classification, represents the number of classification sets, represents the number of subsets corresponding to the classification set, represents the m-th classification set, represents the number of subsets corresponding to the m-th classification set, represents the number of subsets in the m-th classification set, represents the maximum number of subsets in all classification sets under the current classification criteria, represents the minimum number of subsets in all classification sets under the current classification criteria, respectively represent the preset time required for secondary classification and the weight of the degree of uniformity in the classification set, .
[0105] When retrieving the classification criteria stored in the database to perform secondary classification on the preferred data set and obtaining the classification sets after classification, where the preferred data set is a subset of the classification set, and obtaining the two classification sets with the largest and smallest number of subsets, the number of subsets in the two classification sets are respectively denoted as 、 ; when ≥ the preset uniformity threshold, a classification unqualified instruction is issued and secondary classification is performed again; otherwise, a classification qualified instruction is issued and the classification criteria corresponding to this secondary classification are recorded.
[0106] The formula on which the scoring coefficient of the classification criteria is based is as follows:
[0107] (2.2)
[0108] In formula (2.2), represents the scoring coefficient under the current classification criteria, Indicates the number of selections of the y-th classification criterion. Indicates the data classification evaluation value of the y-th classification criterion. Indicates the weights of the preset number of selections and the data classification evaluation value. , according to the scoring coefficient of the classification criterion Select the corresponding classification criteria in descending order and issue them as optimization instructions.
[0109] Analysis of the optimization effect of the instruction
[0110] By using the proportion of the preferred data set in the preliminary retrieval data set, the estimated influence index, and the number of subsets of the preferred data set before and after the instruction optimization during the process of sending "questions" to the "input end", obtain the model optimization efficiency evaluation coefficient for intuitively evaluating the improvement of the efficiency of this retrieval optimization method, so as to have a clear understanding of the entire optimization plan, provide a basis for orderly improvement and adjustment, and then obtain the optimization plan that meets the optimization requirements through the correction of the correction strategy.
[0111]
[0112] In the formula, Indicates the model optimization efficiency evaluation coefficient, Y represents the total number of preferred data sets before the optimization instruction, Indicates the k-th preferred data set before the optimization instruction, Indicates the number of subsets corresponding to the k-th preferred data set before the optimization instruction, X represents the total number of preferred data sets after the optimization instruction, Indicates the x-th preferred data set after the optimization instruction, Indicates the number of subsets corresponding to the x-th preferred data set before the optimization instruction, Respectively represent the weights of the preset proportion of the preferred data set in the preliminary retrieval data set, the estimated influence index, and the ratio of the number of subsets of the preferred data set before and after the instruction optimization, , and when the total number of the preliminary retrieval data set ≤ the average viewing volume, It is recorded as 1. Among them, the historical retrieval trend indicates the number of selections of various classification criteria under the same retrieval terms.
[0113] Correction strategy: It is enabled when the model optimization efficiency evaluation coefficient is lower than the optimization effect threshold. Select the corresponding classification criteria in descending order according to the scoring coefficient, and calculate the corresponding model optimization efficiency evaluation coefficients respectively. Then compare them with the optimization effect threshold, and select the classification criteria corresponding to when the model optimization efficiency evaluation coefficient is lower than the preset optimization effect threshold as the optimization plan output.
[0114] The weight coefficients are determined by the coefficient of variation method. The coefficient of variation method is a method of assigning weights to each evaluation index according to the degree of variation between the current value and the target value of each evaluation index. If the values of a certain index vary greatly and can clearly distinguish each evaluated object, it indicates that the discrimination information of this index is rich, so a larger weight should be given to this index. On the contrary, if the values of each evaluated object on a certain index vary little, then the ability of this index to distinguish each evaluation object is weak, so a smaller weight should be given to this index. This method directly utilizes the information contained in each index and calculates the weights of the indexes, so it has objectivity.
[0115] Embodiment 2: Based on Embodiment 1, this embodiment also provides an optimization method for a large language model based on low-rank adaptation technology. The optimization method includes the following steps:
[0116] Use the input data that has been cleaned and preprocessed as the training dataset for the large language model, train the large language model, and then perform low-rank decomposition and optimization on the initial model by low-rank adaptation technology. Among them, low-rank decomposition is to decompose the original high-dimensional parameter matrix into the product of two or more low-rank matrices;
[0117] Extract the keywords in the received initial instruction as retrieval terms, obtain the retrieval data that can be called by the large language model, and classify them into corresponding retrieval datasets according to types. Then, divide them into a preliminary retrieval dataset and a discardable retrieval dataset according to the number of subsets in the retrieval dataset, and discard the discardable retrieval dataset;
[0118] When receiving the preliminary retrieval dataset, sort it in descending order according to the number of subsets in the preliminary retrieval dataset, select the first several preliminary retrieval datasets as the preferred datasets, and compare the proportion of the preferred datasets in all the preliminary retrieval datasets with the preset screening ratio threshold. When the proportion is less than the screening ratio threshold, then output the preliminary retrieval datasets in the arranged order as the retrieval results; otherwise, output the comparison data as the first retrieval results, analyze and check the estimated impact index of the instruction execution, and issue a check instruction;
[0119] When receiving the check instruction, perform secondary classification on the preferred datasets according to the classification criteria stored in the database, analyze the data classification evaluation value of the secondary classification, combine it with the historical retrieval trend to select the corresponding classification criteria as the optimization instruction output, and then record the instruction after the feedback of the optimization instruction as the secondary instruction, and screen and output the retrieved data after the secondary classification;
[0120] Obtain the proportion of the preferred datasets in the preliminary retrieval datasets during the retrieval process, the estimated impact index, and the number of subsets of the preferred datasets before and after the instruction optimization, generate a model optimization efficiency evaluation coefficient, and when the model optimization efficiency evaluation coefficient is lower than the preset optimization effect threshold, retrieve the optimization scheme correction strategy.
[0121] In the application, several formulas involved are calculated by taking their numerical values after dimensionless treatment. The establishment of the formulas is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. Some coefficients or weights in the formulas are set by those skilled in the art according to the actual situation, so no more details will be given here.
[0122] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0123] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. A large language model optimization system based on low-rank adaptive technology, characterized in that: The optimization system includes: Instruction preliminary analysis module: extract the initial instruction keywords as search terms, obtain the search data that can be called by the large model, classify them into search data sets, and divide them into preliminary and discardable search data sets according to the number of subsets of the search data sets, and discard the latter; Data comparison module: Receive the preliminary search data set, sort it in descending order by the number of subsets, take the first few as the preferred data set, calculate their proportion in the preliminary search data set and compare it with the screening ratio threshold; and according to the proportion result, take the preliminary search data set or the comparison data as the first search result, analyze and estimate the impact index and issue a verification instruction; Instruction optimization module: after receiving the verification instruction, the preferred data set is secondary classified according to the database classification standard, the evaluation value is analyzed, and the optimization instruction is output in combination with the historical search trend. The instruction after feedback is the secondary instruction, which screens and outputs the secondary classified search data. Among them, the historical search trend indicates the number of times various classification standards are selected under the same search term; The analysis process of the data classification evaluation value of the secondary classification is as follows: Retrieve the database storage classification standard to perform secondary classification on the preferred data set and obtain the classified classification set, where the preferred data set is a subset of the classification set, and obtain the two classification sets with the largest and smallest number of subsets. The number of subsets in the two classification sets is recorded as , ; when If the uniformity is greater than or equal to the preset uniformity threshold, a classification failure instruction is issued and a secondary classification is performed again; otherwise, a classification qualification instruction is issued and the classification standard corresponding to the secondary classification is recorded; When receiving the qualified classification instruction, the number of all classification sets and corresponding subsets after secondary classification is obtained, which are recorded as , , and retrieve the two classification sets with the largest and smallest number of subsets based on the classification standard, where the number of subsets is recorded as , ; By analyzing the formula: Get the data classification evaluation value of the secondary classification , where represents the total number of classification sets after secondary classification, represents the mth classification set, represents the number of subsets corresponding to the mth classification set, Indicates the speed of recognition of the number of unit data sets, represents the number of subsets in the mth classification set, All are weights; The scoring coefficient of the classification standard is obtained by calculation ; Scoring coefficients based on classification criteria Select the corresponding classification criteria in order from large to small and issue them as optimization instructions; Model evaluation module: obtains the proportion of the preferred data set, the estimated impact index, and the number of subsets before and after instruction optimization, generates a model optimization efficiency evaluation coefficient, and calls the optimization solution correction strategy when the coefficient is lower than the optimization effect threshold; Among them, the formula based on the model optimization efficiency evaluation coefficient is as follows: In the formula, represents the model optimization efficiency evaluation coefficient, Y represents the total number of preferred data sets before the optimization instruction, represents the kth preferred data set before the optimization instruction, Indicates the number of subsets corresponding to the kth preferred data set before the optimization instruction, Indicates the total number of preliminary retrieval data sets, represents the i-th preliminary retrieval dataset, represents the number of subsets in the i-th preliminary retrieval data set, represents the estimated impact index, X represents the total number of preferred data sets after the optimization instruction, Indicates the xth preferred data set after the optimization instruction. Indicates the number of subsets corresponding to the xth preferred data set before the optimization instruction. Respectively represent weights; When the total number of preliminary search data sets is ≤ the average number of views, Recorded as 1.
2. The large language model optimization system based on low-rank adaptive technology according to claim 1, characterized in that: Before the execution of the instruction preliminary analysis module, the large language model is trained with the cleaned and preprocessed input data, and the initial model is low-rank decomposed and optimized using low-rank adaptive technology; the input data passes through all training layers, and then the output of the training layer is combined with the input adjusted by the new parameters to generate a new adjusted output.
3. The large language model optimization system based on low-rank adaptive technology according to claim 1, characterized in that: The subset of the search data set means: the search data directly obtained according to the search terms; when classifying the search data set, it is compared with the preset rejection threshold; If the number of subsets in the retrieval data set is less than the preset discarding threshold, it is recorded as a discardable retrieval data set; Otherwise, it is recorded as the preliminary search data set; Among them, when searching according to the secondary instruction, if the number of subsets in the discardable search data set increases, and the ratio of the increase to the number of atomic sets is greater than 20%, a reuse instruction is issued to record the corresponding discardable search data set as the preliminary search data set; otherwise, the corresponding discardable search data set is recorded as the non-search data set under the search term.
4. The large language model optimization system based on low-rank adaptive technology according to claim 3, characterized in that: The discard threshold is adjusted according to the following process: Set the adjustment ratio ; Get the number of discardable search datasets under the current search term and the number of preliminary search datasets, and calculate the ratio of the discardable search datasets to the number of preliminary search datasets , and Compare with the preset discard ratio threshold interval; The discard ratio threshold includes: a discard upper limit threshold and a discard lower limit threshold; when >Discard the upper threshold or When the lower threshold is lower than the discard threshold, an adjustment instruction is issued and the adjustment strategy is executed to calculate the average number of subsets of all retrieved data sets. ; When the average number of subsets >Discard upper threshold× , an upward command is issued to ×The upper discard threshold is used as the adjusted upper discard threshold; when the average number of subsets <Discard lower threshold× , a downward adjustment command is issued to ×The discard lower limit threshold is used as the adjusted discard lower limit threshold; Otherwise, no response will be given.
5. The large language model optimization system based on low-rank adaptive technology according to claim 4, characterized in that: The analysis process of the preferred data set ratio is as follows: Retrieve the average reading volume, current reading volume and current total reading time under the previous n search instructions from the database; The average reading volume is used as the number of data selected from the preliminary search data set, and the selected preliminary search data set is recorded as the preferred data set. When the total number of preliminary search data sets is less than or equal to the average reading volume, an early output instruction is issued, and all preliminary search data sets are output as search results; otherwise, the proportion of the preferred data set in all preliminary search data sets is compared with a preset screening ratio threshold; When an advance output instruction is received, if the current reading volume ≥ the average reading volume and the current total reading time ≤ the preset reading time, a call instruction is issued to retrieve the discardable retrieval data set as an unconventional retrieval data set output; otherwise, a residence time distribution map of each retrieval data set under the current reading volume is generated, and the retrieval data set that does not reach the preset standard time is used as a screening criterion, and the retrieval data sets with the same classification as the screening criterion are removed from the unread retrieval data sets.
6. The large language model optimization system based on low-rank adaptive technology according to claim 5, characterized in that: The analysis process for estimating the impact index is as follows: In the process of retrieving the first L retrievals, the number of discarded retrieval data sets and the number of discarded retrieval data sets converted into preliminary retrieval data sets are recorded as , ; The average conversion ratio is calculated by ; The number of scheduled preliminary retrieval data sets and discardable retrieval data sets and their subsets are respectively denoted as , , , ; Through analysis and calculation, the estimated impact index is obtained ; When the impact index is estimated >When the preset impact risk threshold is reached, a data output instruction is issued, and the preliminary search data set is output as the search result in the sorted order; otherwise, a qualified instruction is issued and the verification instruction is executed.
7. The large language model optimization system based on low-rank adaptive technology according to claim 1, characterized in that: Correction strategy: It is turned on when the model optimization efficiency evaluation coefficient is lower than the optimization effect threshold. The corresponding classification standards are selected in order from large to small according to the scoring coefficients, and the corresponding model optimization efficiency evaluation coefficients are calculated respectively. Then, they are compared with the optimization effect threshold. The classification standard corresponding to the model optimization efficiency evaluation coefficient when it is lower than the preset optimization effect threshold is selected as the optimization solution output.
8. A large language model optimization method based on low-rank adaptive technology, using any system described in claims 1 to 7, characterized in that: The optimization method comprises the following steps: The cleaned and preprocessed input data is used as the data set for training the large language model. The large language model is trained, and then the initial model is low-rank decomposed and optimized by low-rank adaptive technology. The low-rank decomposition is to decompose the original high-dimensional parameter matrix into the product of two or more low-rank matrices. Extract keywords in the received initial instruction as search terms, obtain search data that can be called by the large language model, and divide them into corresponding search data sets according to the type. Then, divide them into preliminary search data sets and discardable search data sets according to the number of subsets in the search data sets, and discard the discardable search data sets; When receiving the preliminary search data set, the subsets in the preliminary search data set are sorted from large to small according to the number of subsets, the first several preliminary search data sets are selected as preferred data sets, and the proportion of the preferred data sets in all preliminary search data sets is compared with the preset screening ratio threshold. When the proportion is less than the screening ratio threshold, the preliminary search data sets are output as search results in the sorted order; otherwise, the comparison data is output as the first search result, and the estimated impact index of the execution of the verification instruction is analyzed, and the verification instruction is issued; When receiving the verification instruction, the preferred data set is secondary classified according to the classification standard stored in the database, and the data classification evaluation value of the secondary classification is analyzed, and the corresponding classification standard is selected as the optimization instruction output in combination with the historical search trend. After that, the instruction after the feedback of the optimization instruction is recorded as the secondary instruction, and the search data after the secondary classification is screened and output; Obtain the proportion of the preferred data set in the preliminary search data set during the retrieval process, the estimated impact index, and the number of subsets of the preferred data set before and after instruction optimization, generate a model optimization efficiency evaluation coefficient, and when the model optimization efficiency evaluation coefficient is lower than the preset optimization effect threshold, call the optimization solution correction strategy.
Citation Information
Patent Citations
Large language model fine tuning method based on low-rank matrix decomposition
CN118153715A
Large language model acceleration method and device
CN118569324A
Liver cancer auxiliary diagnosis and question answering method and system based on large language model and medium
CN116975241A
Industrial pollution knowledge relation extraction method based on enhanced knowledge retrieval and large language model collaborative optimization
CN119312897A