Information processing method and device, electronic equipment and chip
By predicting the fine-tuning results based on the relationship between the fine-tuning data volume and the model performance indicators before the model fine-tuning is fine-tuning, the problem of insufficient prediction of the model fine-tuning results in the prior art is solved, and resource and time waste is reduced to ensure that the model performance meets the needs.
Patent Information
- Application Number
- CN202510127796.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art fails to effectively predict the fine-tuning results during the model fine-tuning process, resulting in a possible wasted computing resources and time.
Before fine-tuning the model, based on the current fine-tuning available data or the performance indicator requirements after fine-tuning the model, the prediction results of model fine-tuning are determined in the relationship between the fine-tuning data amount associated with the target business scenario and the model performance indicators.
It provides a basis for whether to perform the fine-tuning process, ensuring that the fine-tuning model performance can meet user needs, and reducing the waste of resources and time caused by the fine-tuning process.
Smart Images

Figure CN120046734A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of model fine-tuning, and particularly to an information processing method, apparatus, electronic device, and chip. Background Art
[0002] Fine-tuning is of great significance in the fields of machine learning and deep learning, and can well improve product performance in specific scenarios. Fine-tuning can not only improve the model performance. By further training on the data of a specific task, the model can learn more specific features and patterns, thus achieving better performance on this task. Fine-tuning can also reduce the need for labeled data. Since the pre-trained model has learned a large amount of general knowledge, only a small amount of labeled data is required to fine-tune the pre-trained model, and the training speed is faster. Fine-tuning also helps to improve the generalization ability of the model, making it better adapt to new and unseen data. In short, fine-tuning is an important bridge connecting pre-trained models and practical applications, enabling the model to better adapt to specific tasks and environments, thereby improving the performance and practicality of the model. Summary of the Invention
[0003] The present disclosure aims to at least solve one of the technical problems in the related art to some extent.
[0004] An embodiment of the first aspect of the present disclosure provides an information processing method, including:
[0005] Receiving a model fine-tuning result prediction request, where the prediction request includes a target service identifier associated with a target model and reference parameters, and the reference parameters are any one of the following: the target performance metric of the target model, the amount of data available during the fine-tuning process;
[0006] Obtaining model description information associated with the target service identifier;
[0007] Determining a prediction result corresponding to the reference parameter based on the model description information;
[0008] Returning the prediction result.
[0009] An embodiment of the second aspect of the present disclosure provides an information processing apparatus, including:
[0010] A receiving module, configured to receive a model fine-tuning result prediction request, where the prediction request includes a target service identifier associated with a target model and reference parameters, and the reference parameters are any one of the following: the target performance metric of the target model, the amount of data available during the fine-tuning process;
[0011] An obtaining module, configured to obtain model description information associated with the target service identifier;
[0012] A prediction module, configured to determine a prediction result corresponding to the reference parameter based on the model description information;
[0013] A sending module, configured to return the prediction result.
[0014] An embodiment of the third aspect of the present disclosure provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the information processing method proposed in the embodiment of the first aspect of the present disclosure is implemented.
[0015] An embodiment of the fourth aspect of the present disclosure provides a chip, which includes a processing circuit and an interface circuit; wherein, the interface circuit is configured to obtain an instruction and send the instruction to the processing circuit, and the processing circuit is configured to execute the instruction to implement the information processing method proposed in the embodiment of the first aspect of the present disclosure.
[0016] An embodiment of the fifth aspect of the present disclosure provides a computer-readable storage medium, storing a computer program, which when executed by a processor, implements the information processing method proposed in the embodiment of the first aspect of the present disclosure.
[0017] The information processing method, device, electronic device, and chip provided by the present disclosure have the following beneficial effects:
[0018] In the embodiment of the present disclosure, before fine-tuning the model, based on the current available data volume for fine-tuning or the performance index requirements after fine-tuning the model, in the relationship between the fine-tuning data volume associated with the target business scenario and the model performance index, the prediction result of the model fine-tuning is determined, thereby providing a basis for whether to perform the fine-tuning process, ensuring that the performance of the fine-tuned model can meet the user's needs, and reducing the waste of resources and time caused by the fine-tuning process.
[0019] Additional aspects and advantages of the present disclosure will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present disclosure. Description of the Drawings
[0020] The above and / or additional aspects and advantages of the present disclosure will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0021] Figure 1 It is a schematic flowchart of an information processing method provided by an embodiment of the present disclosure;
[0022] Figure 2 It is a schematic flowchart of an information processing method provided by another embodiment of the present disclosure;
[0023] Figure 3Schematic diagram for fitting data points associated with a target service identifier provided by an embodiment of the present disclosure;
[0024] Figure 4 Schematic structural diagram of an information processing device provided by another embodiment of the present disclosure;
[0025] Figure 5 Block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure;
[0026] Figure 6 Schematic structural diagram of a chip proposed by an embodiment of the present disclosure. Specific embodiments
[0027] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as limiting the present disclosure.
[0028] In the related art, for model fine-tuning, data associated with actual business or specific tasks is usually directly used to fine-tune a pre-trained model to obtain a fine-tuned model. This fine-tuning method does not predict whether the fine-tuning result meets the requirements, and may waste computing resources because the amount of data is small and the performance of the fine-tuned model still cannot meet the requirements.
[0029] Therefore, the information processing method proposed by the present disclosure can, before fine-tuning, based on the current available data volume for fine-tuning or the performance requirements after model fine-tuning, combined with the relationship between the data volume associated with the current business scenario and the model accuracy, predict the fine-tuning result, thereby providing a basis for whether to perform the fine-tuning process, ensuring that the performance of the fine-tuned model can meet the user's needs, and reducing the waste of resources and time caused by the fine-tuning process.
[0030] It should be noted that the information processing method proposed by the present disclosure can be applied not only to the analysis of model fine-tuning effects, but also to the prediction of the accuracy of large models after a given data volume in various specific business scenarios.
[0031] The information processing method, device, electronic device, and chip of the embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0032] Figure 1 Flow chart of an information processing method provided by an embodiment of the present disclosure.
[0033] It should be noted that the information processing method of this embodiment can be applied to an information processing device. In some possible embodiments, the information processing device can be configured in an electronic device or a chip, so that the electronic device or the chip can execute the function of predicting the data volume required for fine-tuning the prediction model or the upper limit of the performance that can be achieved after fine-tuning.
[0034] As Figure 1 shown, the information processing method may include the following steps:
[0035] Step 101, receive a prediction request for the model fine-tuning result.
[0036] Among them, the prediction request may include a target service identifier associated with the target model and reference parameters.
[0037] Among them, the target model refers to the fine-tuned model, which can be obtained by fine-tuning the pre-trained model using business-related data. The target service identifier is an identifier corresponding to the service scenario to which the data used for fine-tuning belongs, such as user feedback during automobile use, image recognition, etc.
[0038] In the embodiments of the present disclosure, since model fine-tuning is to make the model perform better when solving problems in a specific service scenario, each fine-tuned target model is associated with a suitable target service identifier to clarify which service-related data the target model is fine-tuned with and the usage scenario after fine-tuning. Since the data used for fine-tuning the model is different in different service scenarios, the target models associated with different service identifiers are usually different.
[0039] Among them, the reference parameter is the data used to predict the fine-tuning result, and can be any one of the following: the target performance index of the target model, the data volume available during the fine-tuning process.
[0040] Among them, the target performance index refers to the performance requirement that the fine-tuned target model should meet. For example, if the performance index is the model accuracy rate, then the target performance index of the target model is the size of the target accuracy rate that the fine-tuned target model needs to achieve, or the performance index can also be the recall rate, etc.
[0041] Among them, the data volume available during the fine-tuning process refers to the number of data that can be used for model fine-tuning collected in the scenario corresponding to the target service identifier.
[0042] It should be noted that when the reference parameters included in the received prediction request for the model fine-tuning result are different items, different perspectives of prediction results can be obtained. For example, it can be how much data volume is required for model fine-tuning to meet the target performance index, or it can be the upper limit of the performance that can be achieved by the fine-tuned model using the available data volume.
[0043] In the embodiments of the present disclosure, since the difficulty of model fine-tuning may be different in different business scenarios, and the prediction angles may also be different, before predicting the fine-tuning result of the model, the user needs to clarify the application scenario of the fine-tuned model in the prediction request, that is, the target business identifier associated with the target model, and the reference parameters for prediction. Thus, a prediction request for the model fine-tuning result can be received.
[0044] Step 102, obtain model description information associated with the target business identifier.
[0045] Among them, the model description information is used to describe the relationship between the amount of fine-tuning data and the model performance metrics when the model is fine-tuned using the data associated with each business.
[0046] In the present disclosure, the model description information may be a curve with the amount of data as the abscissa and the performance metric as the ordinate, or may be a table describing the mapping relationship between the performance metric and the amount of data. The present disclosure does not limit this.
[0047] In the embodiments of the present disclosure, in different business scenarios, based on the historical data in the business scenario, the model description information corresponding to the business scenario can be determined and associated with the identifier of the business scenario. Thus, when predicting the result of model fine-tuning, according to the target business identifier included in the prediction request, the target business identifier can be searched in the association relationship between all business identifiers and model description information, so as to obtain the model description information associated with the target business identifier.
[0048] It should be noted that in the embodiments of the present disclosure, the process of determining the model description information associated with the target business identifier may include: first collecting the historical data corresponding to the target business identifier as the fine-tuning data. Then, the fine-tuning data is manually labeled, screened, and grouped to obtain a training data set and a test data set. After that, the hyperparameters of model fine-tuning and the fine-tuned models corresponding to different training data sets are determined using the training data set. Then, using the fine-tuned models corresponding to different training data sets and the test data set, the accuracy rate corresponding to each fine-tuned model is determined. Thus, based on the amount of data in the training data set of the fine-tuned model and the accuracy rate corresponding to the model, the model description information associated with the target business identifier can be obtained by fitting or other means.
[0049] Step 103, determine the prediction result corresponding to the reference parameter based on the model description information.
[0050] In the embodiments of the present disclosure, the data amount or performance metric corresponding to the reference parameter can be queried in the model description information to obtain the prediction result corresponding to the reference parameter.
[0051] It can be understood that the model description information describes the relationship between the amount of data used for model fine-tuning and the performance metrics of the fine-tuned model. Under the prediction condition that the reference parameter is one of them, the prediction result should be the value of the other item. That is to say, for different reference parameters of content, prediction results from different perspectives can be obtained.
[0052] Optionally, in response to the reference parameter being the target performance metric, based on the model description information, the target data amount required to achieve the target performance metric can be determined.
[0053] In the embodiments of the present disclosure, after receiving a prediction request, the type of the reference parameter included in the prediction request can be determined, whether it is a performance metric or a data amount. Then, after determining that the reference parameter is the target performance metric, the data amount size corresponding to the value of the target performance metric can be determined in the model description information, so that when it can be determined that the requirements of the target performance metric can be met after model fine-tuning, how much fine-tuning data is required for model fine-tuning, that is, the target data amount.
[0054] Alternatively, in response to the reference parameter being the available data amount, based on the model description information, the first performance metric that can be achieved by the available data amount can be determined.
[0055] In the embodiments of the present disclosure, after determining that the reference parameter is the available data amount, the performance metric size corresponding to the value of the available data amount can be determined in the model description information, so that the upper limit of the performance that the model can reach after fine-tuning the model using the currently collected data amount can be determined, that is, the first performance metric.
[0056] Step 104, return the prediction result.
[0057] In the embodiments of the present disclosure, the obtained prediction result can be returned to the user. Thus, the user can determine whether to continue the model fine-tuning operation according to the prediction result, or determine how much more fine-tuning data is required for model fine-tuning.
[0058] In the embodiments of the present disclosure, before fine-tuning the model, based on the currently available fine-tuning data amount or the performance metric requirements after model fine-tuning, in the relationship between the fine-tuning data amount associated with the target business scenario and the model performance metric, the prediction result of model fine-tuning is determined, thereby providing a basis for whether to perform the fine-tuning process, ensuring that the performance of the fine-tuned model can meet the user's needs, and reducing the waste of resources and time caused by the fine-tuning process.
[0059] Figure 2 It is a schematic flowchart of an information processing method provided by an embodiment of the present disclosure. As Figure 2 shown, the information processing method may include the following steps:
[0060] Step 201, obtain n first data sets, a second data set, and a first model associated with the target service.
[0061] Among them, the first model is a pre-trained model, that is, a model trained on a large-scale data set. The first model has learned a large amount of general knowledge. Therefore, in this disclosure, fine-tuning is performed on the first model, and only a small data set is required to improve the model performance, and less computing resources are required.
[0062] Among them, the first data set is a training data set, and n is an integer greater than 1.
[0063] It should be noted that the first data set is a training data set for fine-tuning the first model. In order to determine the fine-tuning effect of the first model under training data sets with different data volumes, the number n of the first data sets should be an integer greater than 1. The larger the number n, the more accurate the distribution of the fine-tuning effects of the first model under training data sets with different data volumes can be determined. And the data volume of each of the n first data sets should be different.
[0064] It should be noted that the second data set is a test data set for evaluating the performance of each model fine-tuned using the first data set. In this disclosure, by fixing a second data set, the performance of each model fine-tuned using the first data set is predicted respectively, so as to discover the relationship between the fine-tuning data volume and the model performance metrics.
[0065] In the embodiments of this disclosure, the data obtained related to the target service can be grouped to obtain n first data sets and a second data set. Specifically, the data set related to the target service can be first divided into a training data set and a test data set, and it should be ensured that the data volume of the training data set is larger than that of the test data set. The divided test data set is the second data set, and then the training data set is randomly sampled to obtain n first data sets.
[0066] It should be noted that in order to ensure that the data in the data set is not too sparse on a single label and avoid affecting the accuracy and stability of the model, noise data can also be filtered before obtaining the n first data sets and the second data set related to the target service.
[0067] Optionally, a third data set related to the target service can be obtained first.
[0068] Among them, the third data set includes service data and corresponding annotation results.
[0069] In the embodiments of the present disclosure, the time range and acquisition channels of business data may be selected according to the target business, so as to acquire business data within the time range and channels, and obtain a third data set. The third data set contains the data in the first data set and the second data set.
[0070] It should be noted that the annotation results corresponding to the business data in the third data set are manually annotated. The business data can be attributed to relevant labels according to data attributes, and all labels are real labels. The labels correspond to the target business. In the image recognition business, for example, the annotation results corresponding to the data may include labels such as recognition frames, positions, and sizes.
[0071] In the embodiments of the present disclosure, after obtaining the third data set associated with the target business, the business data in the third data set can also be cleaned. For example, when the target business is the summary of user feedback during the use of an automobile, the business data with a text length less than 4 can be deleted, etc.
[0072] Then, in response to the amount of business data associated with any annotation result in the third data set being less than the data volume threshold, it can be determined that any annotation result and the corresponding business data are noise data.
[0073] Among them, the data volume threshold can be an empirical value determined according to experiments, indicating that when the amount of data associated with any annotation result is greater than or equal to this threshold, it can be ensured that there will be no problem of data sparsity on this label, and thus it will not affect the accuracy and stability of the model.
[0074] In the embodiments of the present disclosure, after obtaining the third data set, according to the annotation results corresponding to each business data in the third data set, the business data can be grouped to obtain the amount of business data associated with each annotation result in the third data set. Then, the amount of business data associated with each annotation result can be compared with the set data volume threshold respectively. When it is determined that the amount of business data associated with any annotation result is less than the data volume threshold, it means that using the amount of business data associated with any annotation result to fine-tune the model may affect the accuracy and stability of the model, and then it can be determined that any annotation result and the corresponding business data are noise data.
[0075] It should be noted that the amount of business data associated with each annotation result can also be determined as the occurrence frequency of the label corresponding to the annotation result. Then, the occurrence frequencies of all labels are sorted, and a limited range is selected (for example, the 20 labels with the highest occurrence frequencies), so that the labels and the corresponding business data outside the limited range can be determined as noise data.
[0076] After that, n first data sets and second data sets can be sampled from the other data in the third data set except the noise data.
[0077] In the embodiments of the present disclosure, other data in the third data set except for the noise data can be first divided into a training data set and a test data set, and it should be ensured that the amount of data in the training data set is larger than that in the test data set. The obtained test data set is the second data set, and then the training data set can be randomly sampled to obtain n first data sets.
[0078] In the embodiments of the present disclosure, after obtaining other data in the third data set except for the noise data, the other data can also be constructed into a lightweight data exchange format JSON (JavaScript Object Notation) data set, and the format is a question-and-answer format that meets the input requirements of the large model. The corresponding annotation results of the other data are used as a part of the prompt information to prompt the large model to select from the above annotation results, so as to increase the amount of prompt information and improve the model performance.
[0079] Step 202: Fine-tune the first model based on each first data set to obtain a second model.
[0080] In the embodiments of the present disclosure, the obtained n first data sets can be used to fine-tune the parameters of the first model respectively to obtain a second model after fine-tuning for each first data set.
[0081] It should be noted that hyperparameters are parameters that need to be manually set during the model training process and cannot be directly learned from data. Reasonable configuration of hyperparameters can ensure that the model can maintain or approach the optimal performance state during the fine-tuning process. Therefore, when fine-tuning the model, it is necessary to first fix the fine-tuning hyperparameters.
[0082] Optionally, the difference between the second data volume included in each first data set and a preset value can be determined first.
[0083] Among them, the preset value is the optimal value of the data volume used to determine the hyperparameters set according to the selected fine-tuning method. For example, when the fine-tuning method is the parameter-efficient fine-tuning LoRA (Low-Rank Adaptation of Large Language Models) method, according to the relevant LoRA papers, to ensure the accuracy of hyperparameter analysis, the preset value corresponding to the data volume is 1000.
[0084] In the embodiments of the present disclosure, a preset value may be determined first according to the currently selected method for fine-tuning the model (which may include Adapter fine-tuning, Prefix-Tuning, LoRA, PET, Prompt-Tuning, etc.). Then, the absolute value of the difference between the second data volume included in each first data set and the preset value is calculated to obtain the difference between the second data volume included in each first data set and the preset value.
[0085] Then, a first data set with the smallest corresponding difference can be determined as the reference data set, and based on the reference data set, the third performance metrics corresponding to different hyperparameter values of the first model are determined respectively.
[0086] In the embodiments of the present disclosure, the differences between the second data volume included in each first data set and the preset value can be compared. When the difference corresponding to a certain first data set is the smallest, it can indicate that the data volume included in this first data set is closest to the optimal value of the data volume for determining the hyperparameters. Then, this first data set can be used as the reference data set to analyze and fine-tune the hyperparameters.
[0087] It should be noted that the hyperparameters may include the batch size (batch_size) of model training, the number of times the training data set is traversed (epoch), and the parameters determined according to the selected fine-tuning method, such as the rank of the low-rank matrix in the LoRA method, etc.
[0088] In the embodiments of the present disclosure, the value of each hyperparameter can be adjusted sequentially, and experiments are carried out using the reference data set to determine the third performance metrics corresponding to different hyperparameter values of the first model.
[0089] After that, the target hyperparameter value can be determined based on the third performance metrics corresponding to different hyperparameter groups.
[0090] In the embodiments of the present disclosure, the third performance metrics corresponding to different hyperparameter groups can be compared, and the hyperparameter group corresponding to the highest third performance metric is determined as the target hyperparameter group. When fine-tuning other parameters in the first model later, the target hyperparameter group is fixed so that the fine-tuning of the hyperparameters can meet the improvement of the model performance.
[0091] Then, based on each first data set, other parameters in the first model except the target hyperparameter value are fine-tuned to obtain a second model.
[0092] In the embodiments of the present disclosure, the process of fine-tuning the first model using the first data set is to input the business data in the first data set into the first model, and by adjusting the parameters in the first model, the loss between the result output by the first model and the labeled result corresponding to the business data in the first data set is made as small as possible.
[0093] In the embodiments of the present disclosure, each first data set can fine-tune other parameters in the first model except for the target hyperparameter value to obtain a second model. Therefore, when there are n first data sets, the number of second models obtained after fine-tuning is also n.
[0094] Step 203: Based on the second data set, test the second model to obtain the second performance metric of the second model.
[0095] In the embodiments of the present disclosure, the second data set can be input into each second model respectively. The second model outputs prediction labels for each business data in the second data set, and then based on the prediction labels and the annotation results corresponding to each business data, the accuracy of the second model can be determined, thereby obtaining the second performance metric of the second model.
[0096] Step 204: Based on the first data volume included in each first data set and the corresponding second performance metric, determine the model description information associated with the target business identifier.
[0097] In the embodiments of the present disclosure, the first data volume can be used as the abscissa, and the corresponding second performance metric can be used as the ordinate to determine the point corresponding to each first data set. Then, the curve obtained by fitting all the points corresponding to the first data sets can be determined as the model description information associated with the target business identifier. Alternatively, a mapping relationship can be established between the first data volume included in each first data set and the corresponding second performance metric and stored in a table as the model description information associated with the target business identifier.
[0098] Optionally, based on the first data volume included in the i-th first data set and the corresponding second performance metric, the i-th data point associated with the target business identifier can be determined, where i is an integer less than or equal to n.
[0099] In the embodiments of the present disclosure, the value of the abscissa corresponding to the i-th data point is the first data volume included in the i-th first data set, and the corresponding ordinate is the i-th first data set and the corresponding second performance metric. Since there are n first data sets, n data points associated with the target business identifier can be determined in the above manner.
[0100] Then, based on the positions and / or distribution characteristics of the n data points associated with the target business identifier, the model description information associated with the target business identifier can be determined.
[0101] Among them, the position of the data point can be the abscissa and ordinate corresponding to the data point. The distribution characteristics of the data points can be described by the scatter plot corresponding to the n data points.
[0102] In the embodiments of the present disclosure, the corresponding relationship between the data volume and the model performance can be determined according to the positions of n data points associated with the target service identifier, so as to obtain the model description information associated with the target service identifier. Alternatively, according to the scatter plot corresponding to the n data points, it can be determined what function the distribution of these n data points conforms to, so as to construct a function describing the corresponding relationship between the data volume and the model performance, and obtain the model description information associated with the target service identifier. Alternatively, the position and distribution characteristics can also be used, and nonlinear least squares fitting can be adopted to find the optimal fitting parameters by minimizing the sum of squared residuals, determine the formed fitting curve, and determine the fitting curve as the model description information associated with the target service identifier.
[0103] Optionally, the i-th data point associated with the target service identifier can be determined based on the first data volume included in the i-th first data set and the corresponding second performance index, where i is an integer less than or equal to n.
[0104] Then, the sum of squared residuals between the n data points associated with the target service identifier and each candidate function among the m candidate functions can be determined, where m is an integer greater than 1.
[0105] It should be noted that since the accuracy rate of the model is [0, 1], generally speaking, as the data volume increases, the contribution of the data increment to the model accuracy rate becomes smaller. Therefore, the shape of the candidate function should be closer to 1, and the model curve should be flatter, such as the logit function.
[0106] In the embodiments of the present disclosure, multiple non-linear monotonically increasing functions can be formulated as candidate functions for curve fitting. Then, the residual between the value of the ordinate corresponding to each data point (i.e., the second performance index) and the value corresponding to the abscissa of the data point in a candidate function can be calculated first, and the sum of squared residuals of all data points corresponding to the residuals can be calculated, so that the sum of squared residuals between the n data points associated with the target service identifier and each candidate function can be obtained.
[0107] After that, the model description information associated with the target service identifier can be determined according to the candidate function corresponding to the minimum sum of squared residuals.
[0108] In the embodiments of the present disclosure, the candidate function corresponding to the minimum sum of squared residuals can be determined, which can more accurately reflect the relationship between the fine-tuned data volume and the fine-tuned model performance in the scenario corresponding to the target service identifier. Therefore, the candidate function corresponding to the minimum sum of squared residuals can be determined as the model description information associated with the target service identifier.
[0109] In the embodiments of the present disclosure, by using the fine-tuning data associated with the target service, the relationship between the amount of fine-tuning data and the performance of the fine-tuned model under different services, that is, the model description information, is specifically analyzed, so as to provide data support for predicting the fine-tuning effect of the model under other fine-tuning conditions in the future, and a model performance improvement upper limit prediction scheme with low overhead and small time requirements and feasibility is provided.
[0110] The following Figure 3 illustrates the process of determining the model description information associated with the target service identifier. Figure 3 It is a schematic diagram for fitting the data points associated with the target service identifier.
[0111] In Figure 3 , the data volume and performance metrics corresponding to each data point are shown in Table 1 below. At this time, the performance metric refers to the accuracy rate.
[0112] Table 1
[0113] Data volume <![CDATA 10,445 > <![CDATA 5,232 > <![CDATA 2,621 > <![CDATA 1,320 > <![CDATA 670 > <![CDATA 335 > <![CDATA 0 > Accuracy <![CDATA 53 .9%]]> <![CDATA 54 .7%]]> <![CDATA 52 .26%]]> <![CDATA 47 .74%]]> <![CDATA 47 .40%]]> <![CDATA 43 .7%]]> <![CDATA 39 .7%]]>
[0114] It should be noted that in Figure 3 , the accuracy rate is converted from the percentage format to the decimal form.
[0115] Therefore, from the data volume and accuracy rate corresponding to each group in Table 1, taken as the abscissa and ordinate corresponding to a data point respectively, 7 data points in Figure 3 can be obtained. Then, by performing power-law fitting on these 7 data points, the curve shown in Figure 3 can be obtained. The curve function is shown in the following formula (1). The relationship between the fitted accuracy rate and the model data volume satisfies:
[0116] y = kx a + b (1)
[0117] where k = 0.0220, a = 0.2302, b = 0.3709, y is the accuracy rate, and x is the data volume.
[0118] Thus, the above formula (1) can be used as the model description information associated with the target service identifier.
[0119] To implement the above embodiments, the present disclosure also proposes an information processing device.
[0120] Figure 4 It is a schematic structural diagram of the information processing device provided by the embodiments of the present disclosure.
[0121] As Figure 4 shown, the information processing device 400 may include:
[0122] A receiving module 401, configured to receive a prediction request for model fine-tuning results, where the prediction request includes a target service identifier associated with the target model and reference parameters, and the reference parameters are any one of the following: a target performance metric of the target model, the amount of data available for the fine-tuning process;
[0123] An obtaining module 402, configured to obtain model description information associated with the target service identifier;
[0124] A prediction module 403, configured to determine a prediction result corresponding to the reference parameters based on the model description information;
[0125] A sending module 404, configured to return the prediction result.
[0126] Optionally, the prediction module 403 can be used for any one of the following:
[0127] In response to the reference parameter being the target performance metric, determine the target data volume required to achieve the target performance metric based on the model description information;
[0128] In response to the reference parameter being the available data volume, determine the first performance metric achievable by the available data volume based on the model description information.
[0129] Optionally, the obtaining module 402 can specifically be used for:
[0130] Obtain n first data sets, a second data set, and a first model associated with the target service, where the first model is pre-trained and generated, and n is an integer greater than 1;
[0131] Based on each first data set, fine-tune the first model to obtain a second model;
[0132] Based on the second data set, test the second model to obtain the second performance metric of the second model;
[0133] Based on the first data volume included in each first data set and the corresponding second performance metric, determine the model description information associated with the target service identifier.
[0134] Optionally, the obtaining module 402 can specifically be used for:
[0135] Determine the difference between the second data volume included in each first data set and a preset value;
[0136] Determine the first data set with the smallest corresponding difference as the reference data set;
[0137] Based on the reference data set, respectively determine the third performance metrics corresponding to different hyperparameter values of the first model;
[0138] Determine a target hyperparameter value based on third performance metrics corresponding to different hyperparameter values;
[0139] Based on each first data set, fine-tune other parameters in the first model except the target hyperparameter value to obtain a second model.
[0140] Optionally, the obtaining module 402 can be specifically used for:
[0141] Obtain a third data set associated with the target service, where the third data set includes service data and corresponding annotation results;
[0142] In response to the service data volume associated with any annotation result in the third data set being less than the data volume threshold, determine any annotation result and the corresponding service data as noise data;
[0143] Sample n first data sets and second data sets from other data in the third data set except the noise data.
[0144] Optionally, the obtaining module 402 can be specifically used for:
[0145] Based on the first data volume included in the i-th first data set and the corresponding second performance metric, determine the i-th data point associated with the target service identifier, where i is an integer less than or equal to n;
[0146] Based on the positions and / or distribution characteristics of n data points associated with the target service identifier, determine the model description information associated with the target service identifier.
[0147] Optionally, the obtaining module 402 can be specifically used for:
[0148] Based on the first data volume included in the i-th first data set and the corresponding second performance metric, determine the i-th data point associated with the target service identifier, where i is an integer less than or equal to n;
[0149] Determine the sum of squared residuals between n data points associated with the target service identifier and each candidate function among m candidate functions, where m is an integer greater than 1;
[0150] According to the candidate function corresponding to the minimum sum of squared residuals, determine the model description information associated with the target service identifier.
[0151] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.
[0152] The information processing device according to the embodiments of the present disclosure determines the prediction result of model fine-tuning based on the relationship between the fine-tuning data volume associated with the target business scenario and the model performance metrics before fine-tuning the model, based on the current available data volume for fine-tuning or the requirements for the performance metrics of the model after fine-tuning, thereby providing a basis for whether to perform the fine-tuning process, ensuring that the performance of the fine-tuned model can meet the user's needs, and reducing the waste of resources and time caused by the fine-tuning process.
[0153] To implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the information processing method proposed in the foregoing embodiments of the present disclosure.
[0154] Figure 5 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 5 The shown electronic device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0155] As Figure 5 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0156] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (hereinafter referred to as: ISA) bus, the Micro Channel Architecture (hereinafter referred to as: MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (hereinafter referred to as: VESA) local bus, and the Peripheral Component Interconnection (hereinafter referred to as: PCI) bus.
[0157] The electronic device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0158] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive"). Although Figure 5 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as compact disc read only memory (CD-ROM), digital video disc read only memory (DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data media interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0159] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present disclosure.
[0160] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0161] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.
[0162] To implement the foregoing embodiments, the present disclosure also proposes a chip, which includes a processing circuit and an interface circuit; wherein, the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute the instructions to implement the data storage method proposed in the foregoing embodiments of the present disclosure.
[0163] Figure 6 is a schematic structural diagram of the chip proposed in the embodiments of the present disclosure. Reference can be made to Figure 6 the schematic structural diagram of the chip 600 shown, but not limited thereto.
[0164] The chip 600 includes a processing circuit 601, and the processing circuit 601 is configured to execute any of the above methods.
[0165] In some embodiments, the chip 600 further includes one or more interface circuits 602. Optionally, the interface circuit 602 is connected to the memory 603. The interface circuit 602 can be used to receive signals from the memory 603 or other devices, and the interface circuit 602 can be used to send signals to the memory 603 or other devices. For example, the interface circuit 602 can read the instructions stored in the memory 603 and send the instructions to the processing circuit 601.
[0166] In some embodiments, the interface circuit 602 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 601 performs other steps.
[0167] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. may be used interchangeably.
[0168] In some embodiments, the chip 600 further includes one or more memories 603 for storing instructions. Optionally, all or part of the memories 603 may be outside the chip 600.
[0169] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program, which when executed by a processor, implements the information processing method as proposed in the foregoing embodiments of the present disclosure.
[0170] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0171] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0172] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0173] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0174] It should be understood that the various parts of the present disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0175] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0176] In addition, in various embodiments of the present disclosure, each functional unit may be integrated in a processing module, may exist physically alone for each unit, or two or more units may be integrated in one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0177] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. An information processing method, characterized in that: include: Receiving a prediction request for a model fine-tuning result, wherein the prediction request includes a target business identifier associated with the target model, and a reference parameter, wherein the reference parameter is any one of the following: a target performance indicator of the target model, and an amount of data available for a fine-tuning process; Acquire model description information associated with the target service identifier; Based on the model description information, determining a prediction result corresponding to the reference parameter; Returns the prediction result.
2. The method according to claim 1, characterized in that The determining, based on the model description information, a prediction result corresponding to the reference parameter, includes any of the following: In response to the reference parameter being the target performance indicator, determining a target data volume required to achieve the target performance indicator based on the model description information; In response to the reference parameter being the available data volume, a first performance indicator achievable by the available data volume is determined based on the model description information.
3. The method according to claim 1 or 2, characterized in that The process of determining the model description information associated with the target service identifier includes: Obtain n first data sets, second data sets, and first models associated with the target business, wherein the first model is generated by pre-training, and n is an integer greater than 1; Based on each of the first data sets, fine-tune the first model to obtain a second model; Based on the second data set, testing the second model to obtain a second performance indicator of the second model; Based on the first data volume contained in each of the first data sets and the corresponding second performance indicator, the model description information associated with the target business identifier is determined.
4. The method according to claim 3, characterized in that The step of fine-tuning the first model based on each of the first data sets to obtain a second model comprises: Determine a difference between the amount of second data included in each of the first data sets and a preset value; Determine a first data set with the smallest corresponding difference as a reference data set; Based on the reference data set, respectively determine third performance indicators of the first model corresponding to different hyperparameter values; Determining a target hyperparameter value based on a third performance indicator corresponding to the different hyperparameter values; Based on each of the first data sets, other parameters in the first model except the target hyperparameter value are fine-tuned to obtain a second model.
5. The method according to claim 3, characterized in that The obtaining of n first data sets and second data sets associated with the target business includes: Acquire a third data set associated with the target business, wherein the third data set includes business data and corresponding annotation results; In response to the business data volume associated with any annotation result in the third data set being less than a data volume threshold, determining that any annotation result and the corresponding business data are noise data; The n first data sets and the second data sets are obtained by sampling the data other than the noise data in the third data set.
6. The method according to claim 3, characterized in that The determining, based on the first data volume contained in each of the first data sets and the corresponding second performance indicator, the model description information associated with the target service identifier includes: Determine, based on the first data volume contained in the i-th first data set and the corresponding second performance indicator, the i-th data point associated with the target service identifier, where i is an integer less than or equal to n; Based on the positions and / or distribution characteristics of the n data points associated with the target business identifier, the model description information associated with the target business identifier is determined.
7. The method according to claim 3, characterized in that The determining, based on the first data volume contained in each of the first data sets and the corresponding second performance indicator, the model description information associated with the target service identifier includes: Determine, based on the first data volume contained in the i-th first data set and the corresponding second performance indicator, the i-th data point associated with the target service identifier, where i is an integer less than or equal to n; Determine the residual sum of squares between n data points associated with the target service identifier and each of the m candidate functions, where m is an integer greater than 1; The model description information associated with the target service identifier is determined according to the candidate function corresponding to the minimum residual square sum.
8. An information processing device, characterized in that: include: A receiving module, configured to receive a prediction request for a model fine-tuning result, wherein the prediction request includes a target business identifier associated with the target model, and a reference parameter, wherein the reference parameter is any one of the following: a target performance indicator of the target model, and an amount of data available for the fine-tuning process; An acquisition module, used to acquire model description information associated with the target business identifier; A prediction module, used to determine a prediction result corresponding to the reference parameter based on the model description information; The sending module is used to return the prediction result.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the information processing method according to any one of claims 1 to 7 is implemented.
10. A chip, characterized in that: The chip includes a processing circuit and an interface circuit; wherein the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute the instructions to implement the information processing method as described in any one of claims 1-7.
11. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the information processing method according to any one of claims 1 to 7 is implemented.