An optimization method for large model systems
By optimizing reference counting and pre-training corpus in the inference process of large models, the problem that large models are difficult to ensure service quality when resource costs are limited is solved, and the effect of reducing the requirements of computing and storage resources is achieved, while ensuring the performance and accuracy of the model.
Patent Information
- Application Number
- CN202410092904.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-01-22
AI Technical Summary
When large models are limited in resource costs, it is difficult to ensure the quality of inference service. Existing model compression methods may lead to information loss, affecting model performance and accuracy.
By counting the model data involved in the feedforward neural network calculation during the model inference process, the low-frequency reference data is determined, and the low-quality corpus is eliminated during the pre-training process, the model training corpus is optimized, the model parameters are reduced, and the computing and storage resource requirements are reduced.
While reducing the cost of model data storage and computing, the quality of the model's inference service is guaranteed, the professionalism of the training corpus is improved, and the computing and storage resources required in the model training and inference process are reduced.
Smart Images

Figure CN117910582B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular, relates to an optimization method for large model systems. Background Art
[0002] Large models contain a large amount of model data, which requires a large amount of storage and computing resources during the training and reasoning process, increasing the cost of computing and storage resources.
[0003] Currently, model compression methods are mainly used to reduce the computing and storage costs of large models. This method uses various compression technologies such as pruning technology, quantization technology, knowledge distillation technology, or sparse activation technology to reduce the size of the model, thereby reducing storage and computing requirements. However, due to the use of compression technology, this method may cause information loss, lead to a decrease in data accuracy, and affect the performance and accuracy of the model.
[0004] Therefore, it is necessary to provide an optimization method for large model systems that can reduce the model data storage cost and computing cost while ensuring the model's reasoning service quality. Summary of the invention
[0005] The embodiments of the present application provide an optimization method for large model systems, which can solve the problem that it is difficult to ensure the quality of reasoning services when large model applications are limited by resource costs.
[0006] The present application embodiment provides an optimization method for a large model system, including:
[0007] During the target model inference process, reference counting is performed on the participating calculation data in the target model to obtain a reference count value, wherein the participating calculation data is model data participating in the feedforward neural network calculation in the target model;
[0008] Determine the model data whose reference count value is less than a first preset value as low-frequency reference data;
[0009] Identify low-quality corpora based on low-frequency citation data;
[0010] The low-quality corpus in the pre-training corpus is eliminated to obtain the training corpus, which is used to train the optimized model.
[0011] The above method counts the references of the model data involved in the feedforward neural network calculation during the model reasoning process to determine the actual use of the model; and adds a pre-training process before model training, through which the training corpus is screened, low-quality corpus is eliminated, the professionalism of the training corpus is improved, and the performance of the optimized model obtained by formal training is guaranteed. In addition, the above method eliminates low-quality corpus according to the application of the data involved in the calculation, that is, the model is optimized according to the actual use of the model, which reduces the parameters that the trained model will generate, and increases the probability of the emergence of lower computing power, thereby reducing the time occupied by training for computing power, reducing the computing and storage resources required in model training and subsequent reasoning, and avoiding the optimization of key data, thereby ensuring the service quality of the model.
[0012] In a possible implementation, the step of determining low-quality corpus according to low-frequency citation data includes:
[0013] In the process of training to obtain the target model, the corresponding relationship between the training corpus used in the training and the model data in the target model is recorded to obtain record information;
[0014] Correspondingly, the steps of determining low-quality corpus based on low-frequency citation data include:
[0015] According to the low-frequency citation data, query the record information to obtain the corpus corresponding to the low-frequency citation data;
[0016] The low-quality corpus is determined based on the corpus corresponding to the low-frequency citation data.
[0017] In the process of training the target model, the above method records the training corpus and the corresponding model data, so as to facilitate reverse query of the corpus corresponding to the low-frequency citation data after determining the low-frequency citation data, determine the low-quality corpus, and also facilitate determination of all model data corresponding to a specified corpus.
[0018] In a possible implementation, the step of determining low-quality corpus according to the corpus corresponding to the low-frequency citation data includes:
[0019] Calculate the similarity between the pre-training corpus and the corpus corresponding to the low-frequency citation data to obtain the corpus similarity;
[0020] The corpus whose corpus similarity is greater than the second preset value is determined as low-quality corpus.
[0021] The above method calculates the similarity between the pre-training corpus and the corpus corresponding to the low-frequency citation data, determines the low-quality corpus in the pre-training corpus, screens and removes the low-quality corpus in the pre-training corpus, and improves the professionalism of the training corpus.
[0022] In a possible implementation, the step of determining low-quality corpus based on low-frequency citation data includes:
[0023] According to the record information, it is determined that among the multiple model data corresponding to the specified corpus, model data exceeding a preset proportion are low-frequency reference data, and the specified corpus is determined as low-quality corpus.
[0024] The above method determines whether the training corpus is low-quality corpus according to the proportion of low-frequency citation data corresponding to the pre-training corpus, thereby avoiding the pre-training corpus being eliminated as low-quality corpus when the pre-training corpus only corresponds to a small amount of low-frequency citation data, thereby reducing the data volume of the final training corpus.
[0025] In a possible implementation, the above optimization method for a large model system further includes:
[0026] During the target model reasoning process, model data whose reference count value is less than a first preset value is stored as an object.
[0027] The above method further reduces the cost of data storage by storing low-frequency reference data in object storage, taking advantage of the fact that object storage does not require the maintenance of complex directory and file structures, and does not require the purchase of expensive storage devices.
[0028] In a possible implementation, after the step of storing the model data whose reference count value is less than the first preset value as an object, the step further includes:
[0029] When the storage time of the model data object whose reference count value is less than the first preset value is greater than the first preset time, the model data whose reference count value is less than the first preset value is discarded.
[0030] The above method discards data when the data has a low reference frequency for a long time, optimizes the model, and processes redundant data in the large model, thereby further reducing storage costs without affecting the precision and accuracy of the model.
[0031] In a possible implementation, it is characterized in that the above optimization method for a large model system further includes:
[0032] During the target model reasoning process, the calculation data involved in the reasoning result generation process and the judgment information of the reasoning result are recorded to obtain the reasoning record information, and the judgment information is made by the user on the quality of the reasoning result;
[0033] Count the data involved in the calculation according to the inference record information to obtain the reference count value of the data involved in the calculation;
[0034] The weight and / or bias parameter of the data involved in the calculation is determined according to the reference count value of the data involved in the calculation after the second preset time length and the judgment information. The second preset time length should be shorter than the first preset time length.
[0035] The above method determines the weights corresponding to the model parameters according to the number of times the model data is used in the target model reasoning process and the quality of the corresponding result judgment, so that the judgment of the generated result has statistical characteristics and the weight is more reasonable. The statistical data is used to optimize the weight algorithm, so that the model reasoning results are more in line with human cognition and safety needs.
[0036] In a possible implementation, the step of determining the weight and / or bias parameter of the participating calculation data according to the reference count value of the participating calculation data after the second preset time period and the determination information includes:
[0037] Obtaining original weights and / or original bias parameters of the data involved in the calculation;
[0038] When the reference count value of the participating calculation data after the second preset time period is greater than the third preset value and the inference result is determined to be good, the value of the original weight and / or the original bias parameter is increased to obtain the weight and / or bias parameter of the participating calculation data;
[0039] When the reference count value of the data involved in the calculation after the second preset time period is less than the third preset value or the inference result is judged to be bad, the value of the original weight and / or original bias parameter is reduced to obtain the weight and / or bias parameter of the data involved in the calculation.
[0040] The above method adopts memory enhancement for model data with high citation frequency and forgetting learning for vectors with no citation or low citation frequency, thereby improving the weight / bias parameters of model data with large contribution and making the final weight more reasonable.
[0041] In a possible implementation, the reasoning record information further includes a record prompt word, and the above optimization method for a large model system further includes:
[0042] Counting the recorded prompt words according to the inference recorded information to obtain the number of prompts of the recorded prompt words;
[0043] When the number of prompts of the recorded prompt word is greater than a fourth preset value, the inference result corresponding to the recorded prompt word is saved;
[0044] Get the user input prompt word;
[0045] When it is determined that the similarity between the prompt word input by the user and the recorded prompt word is greater than a fifth preset value, the inference result corresponding to the recorded prompt word is output as the generated result.
[0046] The above method saves the inference results of prompt words with high prompt frequency, so that when similar prompt words appear, the generation results can be directly obtained from the cached data, which reduces the consumption of computing resources in the inference process and improves the generation speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 It is a flow chart of an optimization method for a large model system provided in an embodiment of the present application;
[0049] Figure 2 It is a schematic diagram of the encoding and decoding process of the large model;
[0050] Figure 3 It is a schematic diagram of the reference counting process provided by an embodiment of the present application;
[0051] Figure 4 It is a schematic diagram of the classification and marking process provided in an embodiment of the present application;
[0052] Figure 5 This is a schematic diagram of a low-quality corpus elimination process provided by an embodiment of the present application;
[0053] Figure 6 This is a schematic diagram of the relationship recording process between corpus and participating computing data provided by an embodiment of the present application;
[0054] Figure 7 It is a schematic diagram of a process of marking low-quality corpus provided by an embodiment of the present application;
[0055] Figure 8 This is a schematic diagram of the classification storage process provided by a specific embodiment of the present application;
[0056] Fig. 9 It is a flowchart of an optimization method for a large model system provided by an embodiment of the present application;
[0057] Fig.10 It is a schematic diagram of the reasoning process provided by an embodiment of the present application. DETAILED DESCRIPTION
[0058] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0059] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0060] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0061] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0062] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0063] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0064] In order to reduce the application cost of large models, it is necessary to reduce the cost of computing and storage resources. The current solutions mainly include compression models and methods of training basic large models with professional corpora. Among them, the compression model reduces redundancy and resource usage through methods such as distillation, quantization, and pruning. However, while the compression model method reduces the application cost, it also reduces the service quality. For example, the quantization process will cause information loss. For some comprehension tasks, the quantized model may not be able to accurately understand and express, and it is difficult to produce accurate and natural answers. For the method of using professional corpora to train basic large models, the large model obtained by this method is of high professional level, but it is not based on a general large model, and the processing of natural language may not be in place. The final generated response may have logical problems, resulting in reduced service quality.
[0065] The present application found that during the training process of the large model, the training corpus is mainly crawled from the Internet, including both low-quality and high-quality corpora. For example, the corpus of GPT-3 (Generative Pre-trained Transformer) includes 67% of regular web crawled data sets, 15% of clean crawled data sets (such as ad-free), and 18% of high-quality training corpus data sets from Github, Wikipedia, papers, stock trading information, etc. Therefore, the present application proposes an optimization method for large model systems, which screens the pre-training corpus through a pre-training process to reduce model parameters, reduce the cost of computing and storage resources required by the model during the inference process, and enable the trained model to have a generation effect similar to that of a large parameter model.
[0066] See also Figure 1 , the embodiment of the present application proposes an optimization method for a large model system, including:
[0067] Step S102: during the target model inference process, reference counting is performed on the participating calculation data in the target model to obtain a reference count value, wherein the participating calculation data is model data participating in the feedforward neural network calculation in the target model;
[0068] Step S104: determining the model data whose reference count value is less than a first preset value as low-frequency reference data;
[0069] Step S106: determining low-quality corpus according to low-frequency citation data;
[0070] Step S108: Eliminate low-quality corpora from the pre-training corpora to obtain training corpora, which are used to train the optimized model.
[0071] Generally, when a model is trained, its training corpus includes both low-quality and high-quality corpus, and may include corpus from multiple application scenarios. However, as a product, the model may only involve a single application field during use, resulting in a large amount of idle data in the model, which puts pressure on the model data storage and calculation.
[0072] In this embodiment, model data, i.e., model parameters, are a collection of tensor data related to the model structure. For example, a transformer model using a decoder-only framework is composed of multiple identical layers, each of which includes a self-attention block and a multi-layer perceptron (MLP) block. The model parameters of the self-attention block include the weight matrix and bias of Q, K, and V, and the weight matrix and bias of output O; the multi-layer perceptron block is composed of two linear layers, including the weight matrices and biases of these linear layers. The self-attention block and the multi-layer perceptron block each have a layer normalization, which contains two trainable model parameters. In addition, the word embedding matrix also has a large number of parameters.
[0073] The data involved in the calculation includes the data calculated by the feedforward neural network (FFN), namely the prompt word, the parameters of the feedforward neural network, and the generated result. In the Transformer architecture, MLP is actually FFN. The parameters of the feedforward neural network here are the weight matrix and bias of the MLP linear layer; the prompt word is in the form of vector data as the input of the feedforward neural network; the generated result is in the form of vector data as the output of the feedforward neural network.
[0074] Both the target model and the optimized model belong to the big model. It is a model with a huge number of parameters, generally with a parameter number of more than one billion, including but not limited to large language models, general large models, vertical large models, AIGC (AI-Generated Content) and more specific implementations of large models in the future. The big model is not limited to a single model file, it can be a collection of multiple model files, which together provide inference application services to the outside world. For example, the inference program in the big model system loads multiple model files and integrates the big model data to provide services to the outside world; the training program in the big model system loads multiple model files through multiple computing nodes or NPUs (Neural Processing Units), or splits a single model file into multiple model files, calculates in parallel, and aggregates weight data to save as a big model.
[0075] join Figure 2, large models, especially transformer (a deep learning network based on attention mechanism) large models, include FFN (Feed Forward Neural Network) calculations in both encoder and decoder in their model architecture, and the vector data of each calculation result describes the process of prompting the prediction result. Therefore, the actual application of model data can be determined based on the number of times the model data participates in FFN calculations.
[0076] Reference counting refers to counting the number of times the data involved in the calculation in the model is referenced. When the data is accessed, the count is performed once, and each time it is used, it increases by 1. In an optional implementation, the data can be recorded when it is accessed, and the number of times the data is accessed can be determined, thereby determining the reference count value of the data.
[0077] Alternatively, after the model data participates in the FFN calculation, the data participating in the FFN calculation and the result data obtained by the calculation are recorded. One data can be recorded multiple times, and finally a reference count value is obtained according to the number of records. For example, see Figure 3 ,When FFN vector data n is referenced, after FFN calculation, the number of references of FFN vector data n changes from Cn to Cn+1, while the reference counts of other vector data remain unchanged.
[0078] The effect of the FFN parameter matrix is to convert the original coordinate space of the data from linearly inseparable to linearly separable. Optionally, after the data participates in the FFN calculation of the reasoning process, it is recorded in the vector database, and the prompt result process is tracked and recorded through the vector database. According to the recorded information, the query and comparison operations in the vector space are performed using vector similarity retrieval and vector indexing to determine the reference count value of the same data participating in the calculation. By using the vector database to record the model data participating in the FFN calculation, the recording process of prompting the result is increased, and the storage resources are optimized according to the record. On the one hand, it meets the recording needs of vector data, and on the other hand, the vector similarity retrieval of the data recorded therein can be performed through the vector database to facilitate the determination of the number of times the vector is actually referenced.
[0079] For step S104, the first preset value may be set by the model designer. Figure 4, the data participating in the calculation can be classified according to the reference count value of the data participating in the calculation, and the data participating in the calculation can be marked as Hot Storage, Normal Storage, and Cold Storage according to the classification, where ColdStorage is model data whose reference count value satisfies a first preset value, and the low-frequency reference data includes model data marked as ColdStorage and model data not participating in the calculation.
[0080] Regarding step S106, in an optional implementation, the step of determining low-quality corpus according to the corpus corresponding to the low-frequency citation data includes:
[0081] Calculate the similarity between the pre-training corpus and the corpus corresponding to the low-frequency citation data to obtain the corpus similarity;
[0082] The corpus whose corpus similarity is greater than the second preset value is determined as low-quality corpus.
[0083] The second preset value may be set by a model designer.
[0084] This implementation calculates the similarity between the pre-training corpus and the corpus corresponding to the low-frequency citation data, determines the low-quality corpus in the pre-training corpus, and screens out the low-quality corpus in the pre-training corpus, so that the final training corpus data set is smaller and more suitable for scenario applications. The model trained in this way will be able to achieve higher model performance with lower computing power.
[0085] In an optional embodiment, the step of determining low-quality corpus according to low-frequency citation data includes:
[0086] In the process of training to obtain the target model, the corresponding relationship between the training corpus used in the training and the model data in the target model is recorded to obtain record information;
[0087] Correspondingly, the steps of determining low-quality corpus based on low-frequency citation data include:
[0088] According to the low-frequency citation data, query the record information to obtain the corpus corresponding to the low-frequency citation data;
[0089] The low-quality corpus is determined based on the corpus corresponding to the low-frequency citation data.
[0090] For example, see Figure 5 and Figure 6, the corpus data is pre-trained to realize feature extraction, and a pre-trained large model is obtained. The present application adds a corpus-FFN vector relationship record database in the model training process, records the relationship between the corpus and the data involved in the calculation, and obtains pre-training record information. After determining the low-frequency citation data within the first preset time length, the corpus corresponding to the low-frequency citation data is obtained by reverse querying the database, and a quality label is added to the corpus corresponding to the low-frequency citation data, and the corpus with the quality label is similarly calculated with the pre-training corpus to determine the low-quality corpus. After eliminating the low-quality corpus in the pre-training corpus, the training corpus is obtained. Among them, the "corpus-FFN vector relationship record database" with quality labels can reflect the frequency of the corpus in practical applications.
[0091] In the process of training the target model, the above implementation method records the training corpus and the corresponding model data, so as to facilitate reverse query of the corpus corresponding to the low-frequency citation data after determining the low-frequency citation data, determine the low-quality corpus, and also facilitate determination of all model data corresponding to a specified corpus.
[0092] By optimizing the training corpus during the pre-training phase, the negative impact on the model is reduced, and the corpus dataset is automatically optimized. The trained large model has fewer parameters, is more efficient, and is more professional, which will miniaturize the model's operating environment, such as a mobile phone environment.
[0093] Text classification is one of the most common and basic tasks in natural language processing (NLP). At present, the classification of corpus datasets is generally completed in the fine-tuning stage after pre-training, that is, supervised learning of the model using labeled data. These labeled data are manually labeled by data labeling outsourcing companies, or simple programs are used instead of manual labeling. Such labeling methods usually have some human factors, which not only increases the cost of training models, but also does not truly reflect the application of model data.
[0094] In an optional implementation, the step of determining low-quality corpus based on low-frequency citation data includes:
[0095] According to the record information, it is determined that among the multiple model data corresponding to the specified corpus, model data exceeding a preset proportion are low-frequency reference data, and the specified corpus is determined as low-quality corpus.
[0096] For example, see Figure 7, determine that the corpus corresponding to the low-frequency cited data or the data not involved in the calculation is corpus m, and determine that it corresponds to m FFN vectors based on the pre-training record information. If 80% of the FFN vectors are marked as Cold Storage or unmarked by the "prompt result process analysis database" after query, that is, the vectors corresponding to corpus m have basically or completely not been cited and have no application value, then the corresponding corpus m will be marked as low-quality corpus.
[0097] The above implementation method determines whether the training corpus is low-quality corpus based on the proportion of low-frequency citation data corresponding to the pre-training corpus, thereby avoiding the pre-training corpus being eliminated as low-quality corpus when the pre-training corpus only corresponds to a small amount of low-frequency citation data, thereby reducing the data volume of the final training corpus.
[0098] For step S108, in an optional implementation, low-quality corpus is periodically eliminated and pre-training data is continuously screened and optimized. After each elimination, a model with smaller parameters and greater generation effect will be generated, thereby increasing the probability of emergence of lower computing power and reducing the computing time occupied by training.
[0099] The beneficial effects of this embodiment are:
[0100] This embodiment determines the actual use of the model by counting the references of the model data involved in the feedforward neural network calculation during the model reasoning process; and adds a pre-training process before model training, through which the training corpus is screened, low-quality corpus is eliminated, the professionalism of the training corpus is improved, and the performance of the optimized model obtained after training is guaranteed. In addition, this embodiment eliminates low-quality corpus according to the application of the data involved in the calculation, optimizes the model according to the actual use of the model, reduces the parameters that the trained model will generate, and increases the probability of the emergence of lower computing power, thereby reducing the time occupied by training for computing power, reducing the computing and storage resources required in model training and subsequent reasoning, and avoiding the optimization of key data to ensure the service quality of the model.
[0101] According to the above embodiment, in yet another embodiment:
[0102] During the target model reasoning process, model data whose reference count value is less than a first preset value is stored as an object.
[0103] When the storage time of the model data object whose reference count value is less than the first preset value is greater than the first preset time, the model data whose reference count value is less than the first preset value is discarded.
[0104] Optionally, the first preset time period is one quarter or one year.
[0105] Since large models involve massive amounts of data, only a small portion of the data may be used in the calculation in a short period of time, but data objects that are not used in the calculation may still be referenced, so the storage cost of this data needs to be considered. Object storage does not require the maintenance of complex directory and file structures, nor does it require the purchase of expensive storage devices.
[0106] The above implementation method evaluates aging processing based on the time threshold, that is, object storage is used and discarded after a period of time. In this process, the data of the large model will be greatly optimized, especially in the situation awareness process for specific industry applications, which can perceive the commonly used prompt words in the industry and the parameter data in the multi-layer perception block (MLP, Multi-Layer Perceptron) of the large model, providing better reasoning service quality. This can reduce storage and maintenance costs while ensuring data integrity and availability.
[0107] According to the above embodiment, in yet another embodiment:
[0108] See also Figure 8 , with a certain period of time as a cycle, the data involved in the calculation is classified according to the number of references of the data involved in the calculation, and the data involved in the calculation is marked as Hot Storage, Normal Storage, and Cold Storage according to the classification, where Cold Storage is the data involved in the calculation whose reference count value meets the preset conditions. For vector data marked as HotStorage, since it is often used, it can be stored in high-performance, low-latency storage resources, while the data marked as NormalStorage and Cold Storage will be stored in storage resources with lower performance and greater latency.
[0109] For data that does not participate in the calculation, that is, unlabeled vector data, it will enter the aging process. The storage method can be object storage, which can process the redundant data of large models and reasonably allocate storage resources. While taking into account the quality of inference services, it effectively reduces the cost of storage resources.
[0110] Optionally, data that has been in the Cold Storage level for a long time may be aged, that is, the data may be discarded after exceeding a first preset time to save storage resources. Figure 4 As shown, for unmarked data and vector data marked as Cold Storage, aging processing will be performed based on the time threshold for evaluation, that is, object storage will be used, and the data will be discarded after exceeding the first preset time.
[0111] According to the above embodiment, in another embodiment, the large model data optimization method provided by the present application further includes:
[0112] In the process of target model reasoning, the calculation data involved in the process of generating the reasoning result and the judgment information of the reasoning result are recorded to obtain the reasoning record information;
[0113] Count the data involved in the calculation according to the inference record information to obtain the reference count value of the data involved in the calculation;
[0114] The weight and / or bias parameter of the participating calculation data is determined according to the reference count value of the participating calculation data after the second preset time period and the determination information.
[0115] Optionally, the second preset time period may be one month or one quarter. The second preset time period should be shorter than the first preset time period.
[0116] At present, the process of adjusting the model weight is that the prompter judges the generated results, and the model adjusts the weights in the model according to the judgment results. However, in this method, the prompter is often the only one who judges the quality of the results. The weights learned by the model feedback are not large enough, and they are random due to different personal preferences, which is prone to safety hazards.
[0117] In this embodiment, the judgment information of the model data and the reasoning result in the reasoning process can be recorded through the database to obtain the reasoning record information. And based on the quality of each generated result in the reasoning record information and the reference count value of the data involved in the calculation, the weight or bias corresponding to the calculation is optimized. The judgment information of the reasoning result includes two categories: positive and negative. Optionally, the more reference count values of the parameter reasoning data, the higher the weight value when the reasoning judgment result of the model data is positive. On the contrary, the lower the reference count value of the parameter reasoning data, the lower the weight value when the reasoning judgment result of the model data is negative.
[0118] In an optional embodiment, the step of determining the weight and / or bias parameter of the participating calculation data according to the reference count value of the participating calculation data after the second preset time period and the determination information includes:
[0119] Obtaining original weights and / or original bias parameters of the data involved in the calculation;
[0120] When the reference count value of the participating calculation data after the second preset time period is greater than the third preset value and the inference result is determined to be good, the value of the original weight and / or the original bias parameter is increased to obtain the weight and / or bias parameter of the participating calculation data;
[0121] When the reference count value of the data involved in the calculation after the second preset time period is less than the third preset value or the inference result is judged to be bad, the value of the original weight and / or original bias parameter is reduced to obtain the weight and / or bias parameter of the data involved in the calculation.
[0122] The third preset value may be set by the model designer or determined according to the reference count values of all the data involved in the calculation. For example, when the second preset duration is determined, the average reference count value of all the data involved in the calculation is a, and a may be used as the preset value.
[0123] For example, Fig. 9 As shown, the FFN vector, i.e., the data involved in the calculation, is recorded through the "prompt result process analysis database" to obtain the reasoning record information, and the participation frequency of the FFN vector is obtained according to the reasoning record information. According to the reference frequency of the vector, memory enhancement is adopted for the vectors with high frequency references, and forgetting learning is adopted for the vectors with no references or low frequency references. Among them, memory enhancement refers to strengthening parameters such as weights and biases, and forgetting learning refers to weakening parameters such as weights and biases. Among them, when the forgotten parameters are small to a certain extent, they will be set to zero during the quantization (precision reduction) process until they are finally optimized by algorithms such as pruning. After this treatment, the RHLF (Reinforcement Learning from Human Feedback) of the model during future retraining will be consistent with human cognition.
[0124] The beneficial effects of this embodiment are:
[0125] By determining the weights corresponding to the model parameters based on the number of times the model data is used in the target model reasoning process and the quality of the corresponding result judgment, the judgment of the generated result has statistical characteristics and the weight is more reasonable. The weight algorithm is optimized using statistical data to make the model reasoning results more in line with human cognition and safety needs.
[0126] According to the above embodiment, in yet another embodiment:
[0127] The reasoning record information also includes record prompt words. The optimization method for the large model system provided by the present application also includes:
[0128] Counting the recorded prompt words according to the inference recorded information to obtain the number of prompts of the recorded prompt words;
[0129] When the number of prompts of the recorded prompt word is greater than a fourth preset value, the inference result corresponding to the recorded prompt word is saved;
[0130] Get the user input prompt word;
[0131] When it is determined that the similarity between the prompt word input by the user and the recorded prompt word is greater than a fifth preset value, the inference result corresponding to the recorded prompt word is output as the generated result.
[0132] For example, see Fig.10When the user inputs a prompt word, the prompt word input by the user is recorded through the "prompt result process analysis database" to obtain the recorded prompt word, and the recorded prompt word is counted through the reasoning record information. A cache system is built according to the frequency of the recorded prompt word. When the number of prompts of the recorded prompt word is greater than the fifth preset value, the corresponding reasoning result is stored in a separate cache data. When the problem of high similarity occurs, that is, when the user input prompt word and the recorded prompt word of the cached reasoning result are highly similar, the generated result can be taken out of the cache first, instead of calling the background reasoning to generate it every time, thereby reducing the time occupied by reasoning for computing power, accelerating generation, and improving model performance.
[0133] Specifically, before inference, the cache system will be queried to see if there is a cache with high similarity to the user input prompt word. If there is, it will be directly retrieved and used for generation. If not, it will be inferred and generated by itself. This method automatically speeds up the inference calculation and improves the performance of the model.
[0134] It should be noted that, in this embodiment, the first preset value can be set by the model designer, and there is no definite relationship between the first preset value, the second preset value, the third preset value, the fourth preset value, and the fifth preset value unless otherwise stated.
[0135] The beneficial effects of this embodiment are:
[0136] The inference results of prompt words with high prompt frequency are saved, so that when similar prompt words appear, the generation results can be directly obtained from the cached data, which reduces the consumption of computing resources in the inference process and improves the generation speed.
[0137] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in one or more computer-readable storage media. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, USB flash drive, mobile hard disk, disk or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0139] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0140] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0141] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0142] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An optimization method for a large model system, characterized in that: include: During the target model inference process, reference counting is performed on the participating calculation data in the target model to obtain a reference count value, wherein the participating calculation data is model data participating in the feedforward neural network calculation in the target model; Determine the model data whose reference count value is less than a first preset value as low-frequency reference data; Determining low-quality corpus according to the low-frequency citation data; Eliminate low-quality corpora from pre-training corpora to obtain training corpora, which are used to train an optimized model; The method further comprises: During the target model reasoning process, the calculation data involved in the reasoning result generation process and the judgment information of the reasoning result are recorded to obtain reasoning record information, wherein the judgment information is made by the user on the quality of the reasoning result; Counting the data involved in the calculation according to the inference record information to obtain a reference count value of the data involved in the calculation; Obtaining original weights and / or original bias parameters of the data involved in the calculation; When the reference count value of the participating calculation data after the second preset time period is greater than the third preset value and the inference result is determined to be good, increase the value of the original weight and / or the original bias parameter to obtain the weight and / or bias parameter of the participating calculation data; When the reference count value of the participating calculation data after the second preset time period is less than a third preset value or the inference result is determined to be bad, reducing the value of the original weight and / or the original bias parameter to obtain the weight and / or bias parameter of the participating calculation data; The reasoning record information also includes a record prompt word, and the method further includes: Counting the record prompt words according to the inference record information to obtain the number of prompts of the record prompt words; When the number of prompts of the recorded prompt word is greater than a fourth preset value, saving the inference result corresponding to the recorded prompt word; Get the user input prompt word; When it is determined that the similarity between the user input prompt word and the record prompt word is greater than a fifth preset value, outputting the inference result corresponding to the record prompt word as a generated result; The step of determining low-quality corpus according to the low-frequency citation data includes: In the process of training to obtain the target model, recording the corresponding relationship between the training corpus used in the training and the model data in the target model to obtain record information; Correspondingly, the step of determining low-quality corpus according to the low-frequency citation data includes: According to the low-frequency citation data, query the record information to obtain the corpus corresponding to the low-frequency citation data; The low-quality corpus is determined according to the corpus corresponding to the low-frequency citation data.
2. The optimization method for a large model system according to claim 1, characterized in that: The step of determining the low-quality corpus according to the corpus corresponding to the low-frequency citation data comprises: Calculate the similarity between the pre-training corpus and the corpus corresponding to the low-frequency citation data to obtain corpus similarity; The corpus whose corpus similarity is greater than a second preset value is determined as the low-quality corpus.
3. The optimization method for large model systems according to claim 1, characterized in that: The step of determining low-quality corpus according to the low-frequency citation data comprises: According to the record information, it is determined that among the multiple model data corresponding to the designated corpus, model data exceeding a preset proportion are the low-frequency cited data, and the designated corpus is determined to be a low-quality corpus.
4. The optimization method for a large model system according to any one of claims 1 to 3, characterized in that: Also includes: During the target model reasoning process, the model data whose reference count value is less than the first preset value is stored as an object.
5. The optimization method for large model systems according to claim 4, characterized in that: After the step of storing the model data whose reference count value is less than the first preset value as an object, the step further includes: When the storage time of the model data object whose reference count value is less than the first preset value is greater than the first preset time, the model data whose reference count value is less than the first preset value is discarded.
Citation Information
Patent Citations
Corpus processing method, device and equipment and computer storage medium
CN116701595A
System and method for data classification using machine learning during archiving
US20180373722A1