A method and system for optimizing data processing time of large models
By recording the correlation between the processing time and the number of characters of the large model, calculating the degree of discreteness, and dynamically selecting the target model, the problem of inconsistent processing time in the multi-model bridging platform is solved and the processing efficiency is improved.
Patent Information
- Application Number
- CN202510990490.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-18
AI Technical Summary
In the multi-model bridging platform, the processing time of the same input data varies significantly between different large models, the correlation between the number of characters and time is inconsistent, and there is a lack of quantitative standards, resulting in low processing efficiency.
By recording the processing time and number of characters of the candidate large model, establishing a time-character-number correlation, calculating the degree of discreteness, dynamically selecting the target model to optimize the processing time, and using random selection or correlation calculation to determine the optimal model.
The processing time of large model data has been optimized, and the processing efficiency has been improved by 30%-50%, adapting to the processing needs of different scenarios.
Smart Images

Figure CN120509493B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology applications, and in particular to a method and system for optimizing data processing time of a large model. Background Art
[0002] Currently, the application of large models in fields such as intelligent customer service, autonomous driving, and medical diagnosis is experiencing explosive growth, and the number of large model vendors in the market has increased dramatically. Each vendor's model has its own distinct focus on processing capabilities and applicable scenarios. To integrate the advantages of various large models and reduce user adaptation costs, multi-model bridging platforms have become a mainstream architecture in the industry. Users can access different large models through a unified interface without having to worry about the underlying differences between the models, significantly improving application convenience. However, in actual operation, this platform faces a core challenge in optimizing processing time. First, the processing time of different large models for the same input data varies significantly. This difference fluctuates with data characteristics (such as the number of characters in text data). For example, long text data may take less time in models that excel at parallel processing, but may take twice as long in models that focus on accuracy verification. This results in significant fluctuations in processing time between different models for the same input data (known as "discrepancy"). On the other hand, there is a strong correlation between processing time and the number of data characters: in most cases, the more characters there are, the more complex the parsing, conversion, and inference steps of the model, and the longer the processing time. However, the correlation between the number of characters and time varies from model to model, and there is a lack of a unified quantitative standard. Therefore, how to quantitatively analyze the correlation between the number of characters and events in large models, and dynamically determine the model selection strategy based on the discrete degree of processing time for the same input data in different models, has become the key to improving the processing efficiency of the multi-model bridging platform. Developing a method that can optimize model selection based on the number of data characters and time fluctuation characteristics is of great significance for shortening data processing delays and improving platform response speed. It is also a core requirement for promoting the in-depth application of large models in scenarios with high real-time requirements. Summary of the Invention
[0003] In view of the above technical problems, the technical solution adopted by the present invention is:
[0004] According to a first aspect of the present invention, a method for optimizing data processing time of a large model is provided, the method comprising the following steps:
[0005] S100: Input m sample input data into n candidate large models respectively to obtain the processing time record set T of the i-th candidate large model. i ={T i1 , T i2 ,……,T ij ,……,T im}, where T ijIt is the processing time of the i-th candidate large model for the j-th sample input data, and the value of j ranges from 1 to m.
[0006] S200, based on T i , obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model.
[0007] S300, based on the processing time record set corresponding to the n candidate large models, obtain the discrete degree DS corresponding to the j-th sample input data j The degree of dispersion is used to characterize the degree of fluctuation in the processing time difference when different large models process the same sample input data.
[0008] S400, traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m}, for the discrete degree DS traversed j , if DS j >DS0, change DS j The number of characters of the corresponding sample input data is used as the target number of characters.
[0009] S500, in response to receiving the target input data to be processed, obtain the number of characters of the target input data; if the number of characters of the target input data is less than the target number of characters, randomly select one large model from the m candidate large models as the target model for processing the target input data; otherwise, input the number of characters of the target input data into the time-character number association formula of each candidate large model to obtain the processing time of each candidate large model for processing the target input data.
[0010] S600: Using the candidate large model with the shortest processing time as the target model for processing the target input data.
[0011] According to a second aspect of the present invention, a system for optimizing data processing time of a large model is provided, the system comprising:
[0012] The processing time acquisition module is used to input m sample input data into n candidate large models respectively, and obtain the processing time record set T of the i-th candidate large model i ={T i1 , T i2 ,……,T ij ,……,T im}, where T ij It is the processing time of the i-th candidate large model for the j-th sample input data, and the value of j ranges from 1 to m.
[0013] Time character number association acquisition module, used for Ti , obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model.
[0014] The discrete degree acquisition module is used to obtain the discrete degree DS corresponding to the jth sample input data based on the processing time record set corresponding to the n candidate large models j The degree of dispersion is used to characterize the degree of fluctuation in the processing time difference when different large models process the same sample input data.
[0015] The target character number acquisition module is used to traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m}, for the discrete degree DS traversed j , if DS j >DS0, change DS j The number of characters of the corresponding sample input data is used as the target number of characters.
[0016] A target model acquisition module is used to obtain the number of characters of the target input data in response to receiving the target input data to be processed. If the number of characters of the target input data is less than the target number of characters, randomly select one large model from the m candidate large models as the target model for processing the target input data; otherwise, input the number of characters of the target input data into the time-character number association formula of each candidate large model to obtain the processing time of each candidate large model for processing the target input data, and use the candidate large model with the minimum processing time as the target model for processing the target input data.
[0017] The present invention has at least the following beneficial effects:
[0018] The embodiment of the present invention provides a method and system for optimizing the data processing time of a large model, comprising: inputting m samples into n candidate large models to obtain a set of model processing time records Ti; based on T i Get the time character number association formula; get the sample discreteness DS based on the processing time record set j ; Traverse the discrete degree set and convert DS j>DS0 corresponds to the number of sample characters as the target number of characters; upon receiving the target input data to be processed, if its number of characters is less than the target number of characters, a candidate model is randomly selected; otherwise, the number of characters is input into the association formula to obtain the processing time of each model, and the model with the smallest processing time is selected. This invention automatically identifies "high volatility scenarios" and "low volatility scenarios" by calculating the time discreteness of different models processing the same sample and setting a target number of characters. In high volatility scenarios, the optimal model is accurately selected based on the target number of characters, while a simplified selection strategy is adopted in low volatility scenarios. This can optimize the processing time of large model data and improve overall processing efficiency by 30%-50%.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A flowchart of a method for optimizing data processing time of a large model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0024] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be performed in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. A process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0025] The embodiment of the present invention provides a method for optimizing the data processing time of a large model, such as Figure 1 As shown, the method may include the following steps:
[0026] S100: Input m sample input data into n candidate large models respectively to obtain the processing time record set T of the i-th candidate large model. i ={T i1 , T i2 ,……,T ij ,……,T im}, where T ij It is the processing time of the i-th candidate large model for the j-th sample input data, and the value of j ranges from 1 to m.
[0027] In the embodiment of the present invention, the m sample input data need to cover different data scales, data characteristics, and business scenarios to ensure that the acquired processing time data is comprehensive and reliable.
[0028] In an embodiment of the present invention, when each sample input data enters the large model for processing, a timing program is started. When the large model completes the processing task of the sample statement, the timing program stops and the time from the start of input to the end of processing is recorded as the processing time of the large model.
[0029] In the embodiment of the present invention, the time of entering the large model can be recorded in the following manner:
[0030] Code insertion marker: In the data transmission code logic, insert a time-recording code when the target data or sample input data begins to be transferred to the big model. For example, in a data transmission script written in Python, use the time module to record the current timestamp, import time, and add start_time=time.time() before the statement that starts data transmission to mark the starting time of data entering the big model.
[0031] Message queues and logging: If you use message queues to manage data transmission, you can log the current time when data is placed in the message queue and awaits retrieval by the large model. For example, if you use RabbitMQ as a message queue, after the producer sends data to the queue, use a logging tool (such as Python's logging module) to record "Data sent to message queue, time: " + str(datetime.now())." When the large model subsequently retrieves data from the message queue, it can use the logged time to determine when the data entered the large model.
[0032] In the embodiment of the present invention, the time for the large model to complete the processing task can be determined by the following method:
[0033] Output signal monitoring: After completing data processing, large models typically output processing results. A monitoring mechanism can be set up in the code that receives the processing results. When the complete processing result data is received, the current time is recorded as the processing completion time. For example, if the large model returns the processing result through an API interface, add end_time = time.time() to the response receiving function of the API call to obtain the time when the large model completed the processing task.
[0034] Task status callback: Some large model services provide a task status query or callback mechanism. For example, when submitting a data processing task, a task ID is assigned. The task status query interface is periodically called to check whether the task is in the "Completed" state. Alternatively, a callback address can be set. When the large model completes processing a task, a completion notification is sent to that address, and the time of receipt is recorded.
[0035] Resource Usage Monitoring: Large models consume a certain amount of computing resources (such as CPU and GPU usage) when processing data. System resource monitoring tools can be used to monitor the resource usage of large model processes in real time. When resource usage drops significantly and stabilizes at a low level for a period of time (e.g., 1-2 seconds), the large model can be considered to have completed its processing task and the time at which this occurred can be recorded. For example, in Linux, use the top or nvidia-smi commands to obtain resource usage information, and use scripts for real-time analysis and assessment.
[0036] In an embodiment of the present invention, the m sample input data may be layered according to the number of characters, and the samples in each layer may be input into n candidate large models in a multi-threaded parallel manner, with the number of characters in the m sample input data increasing in sequence.
[0037] S200, based on T i, obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model.
[0038] Furthermore, S200 may specifically include:
[0039] S201, calculate T ij Z score Z ij and interquartile range (IQR) i , if T i <Q1 i -1.5IQR i , or, T ij >Q3 i +1.5IQR i , and |Z ij |>2, T ij As an outlier, a replacement value is generated using cubic spline interpolation; Q1 i T i The first quartile, Q3 i T i The third quartile of .
[0040] S202, the number of characters C of the input data of the jth sample j First, the Box-Cox transformation is performed to transform its distribution form, and then the Z score standardization is used to eliminate the dimension effect, and finally C j The converted value after normalization.
[0041] S203, select linear regression model T i (C) = β0 + β1 × C + γ as the time character number correlation formula of the i-th candidate large model, T i (C) is the processing time of the i-th candidate large model, C is the number of characters, β0 and β1 are fitting parameters, and γ is the error term.
[0042] S204, based on the m processing times and character numbers corresponding to the i-th candidate large model, solve the time-character-number correlation formula of the i-th candidate large model to obtain the fitting parameters and error terms of the time-character-number correlation formula of the i-th candidate large model, and then obtain the time-character-number correlation formula of the i-th candidate large model.
[0043] In an embodiment of the present invention, the time character number correlation equation of the i-th candidate large model can be solved using the least squares method or the grid search method, preferably the least squares method. Those skilled in the art should understand that the method of solving the time character number correlation equation of the i-th candidate large model using the least squares method or the grid search method can fall within the scope of the prior art.
[0044] S300, based on the processing time record set corresponding to the n candidate large models, obtain the discrete degree DS corresponding to the j-th sample input data j The degree of dispersion is used to characterize the degree of fluctuation in the processing time difference when different large models process the same sample input data.
[0045] Furthermore, S300 specifically includes:
[0046] S310, obtaining the processing time record set T corresponding to the j-th sample input data j ={T 1j , T 2j ,……,T ij ,……,T nj}, that is, to organize the processing time sets of n candidate large models to process the same sample input data.
[0047] S320, get T j The corresponding time difference set △T j ={△T j 12 , △T j 23 ,……,△T j (i-1)i ,……,△T j (n -1)n}, △T j (i-1)i is △T j The i-1th time difference in j (i-1)i =T ij -T (i-1)j .
[0048] By calculating the time difference between adjacent large models in processing the same statement, we can preliminarily reflect the efficiency differences when different large models process the same data.
[0049] S330, obtain △T j The corresponding degree of dispersion , △T jr is △T j The rth time difference in the time interval, r ranges from 1 to n-1, Avg (△T j ) is △T j The average of the n-1 time differences in .
[0050] S400, traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m}, for the discrete degree DS traversed j , if DS j >DS0, change DSj The number of characters of the corresponding sample input data is used as the target number of characters.
[0051] If DS j >DS0, it means that the efficiency of different large models in processing the sample input data varies greatly. j The number of characters in the corresponding sample input data is used as the target character count. This is because when the data volume reaches this character count, the processing efficiency of each candidate large model begins to diverge significantly. Using this as the threshold, in practical applications, for data with less than this character count, randomly selecting a large model for processing has little impact on overall efficiency. However, for data with a character count greater than or equal to this, selecting a large model through subsequent multi-dimensional evaluation can more effectively match data processing requirements and achieve a balance between efficiency and accuracy.
[0052] In this embodiment of the present invention, DS0 is a threshold for setting the degree of discreteness. This threshold can be set based on the actual business requirements for large-scale model processing efficiency consistency. For example, in scenarios where high efficiency and stability are required, DS0 can be set to a small value. Conversely, in scenarios where efficiency and stability are less required, DS0 can be appropriately increased.
[0053] Those skilled in the art should understand that since the number of characters in the m sample input data increases in sequence, during the traversal process, as long as the DS j >DS0, the traversal of the discrete degree set will no longer continue and the traversal program will end.
[0054] S500, in response to receiving the target input data to be processed, obtain the number of characters of the target input data; if the number of characters of the target input data is less than the target number of characters, randomly select one large model from the m candidate large models as the target model for processing the target input data; otherwise, input the number of characters of the target input data into the time-character number association formula of each candidate large model to obtain the processing time of each candidate large model for processing the target input data.
[0055] S600: Using the candidate large model with the shortest processing time as the target model for processing the target input data.
[0056] ) Furthermore, in another embodiment of the present invention, after S400, the following steps are also included:
[0057] S410, based on the m sample input data and the corresponding real results, obtain the problem accuracy correlation formula of each candidate large model, wherein the problem accuracy correlation formula is used to characterize the relationship between the number of problems in the input data, the problem correlation degree and the accuracy of the large model.
[0058] After S500, the following steps are also included:
[0059] S510: Obtain the number of questions and the relevance of questions in the target input data.
[0060] S520: Input the number of questions and the question relevance corresponding to the target input data into the question accuracy correlation formula of each candidate large model to obtain the corresponding accuracy.
[0061] The S600 is replaced by:
[0062] S610 , based on the processing time and accuracy of each candidate large model, selecting a target model from the m candidate large models as a target model for processing the target input data.
[0063] Furthermore, S410 specifically includes:
[0064] S4101, obtain feature information of each sample input data.
[0065] In one embodiment of the present invention, the characteristic information includes the number of questions and the relevance of the questions.
[0066] In another embodiment of the present invention, the characteristic information may include the number of characters, the number of questions, the relevance of questions, the span of fields, the density of topics, etc.
[0067] In the embodiment of the present invention, the question number can be identified by punctuation marks and question words, or by a pre-trained model or a special question detection model.
[0068] Furthermore, in an embodiment of the present invention, the question relevance of the sample input data or the target input data satisfies the following conditions:
[0069] S10, with a sliding window length of g and a sliding step of 1, divide all the problems of the data to be processed into the corresponding problem record set Q g ={Q g1 , Q g2 ,……,Q gh ,……,Q g(z(g))}, Q gh Q g The h-th problem subset in Q, where h ranges from 1 to z(g), and z(g) is g The number of problem subsets in , z(g) = p-g+1. The value of g ranges from 2 to p, where p is the number of problems in the data to be processed, and the data to be processed is the sample input data or the target input data.
[0070] S20, get Q gh The corresponding question relevance S gh , and based on the obtained z (g) question relevance, obtain Q gThe corresponding question relevance S g .
[0071] In the embodiment of the present invention, S gh The following conditions are met:
[0072] , where sim(Q gh u , Q gh v ) is Q gh The semantic similarity between the u-th question and the v-th question in , where u ranges from h to h+g-1 and v ranges from u+1 to h+g-1.
[0073] In an embodiment of the present invention, the semantic similarity between questions may be cosine similarity, which may be obtained based on the semantic embedding vector of the Sentence-BERT model.
[0074] In the embodiment of the present invention, S g The following conditions are met:
[0075] S g =(∑ z(g) h=1 (w h ×S gh )) / (∑ z(g) h=1 w h ), w h For S gh The weight coefficient of .
[0076] In the embodiment of the present invention, w h =e -α(h / z(g)) ,e is a natural number, α is the attenuation factor, and its value can be 0.5.
[0077] S30 , obtaining the problem relevance of the data to be processed based on the problem relevance corresponding to the p problem record sets.
[0078] In the embodiment of the present invention, the problem correlation of the data to be processed satisfies the following conditions: S=∑ p g=2 (w g ×S g N ), S g N For S g The normalized value of w g For S g N The corresponding weight coefficient, ∑ p g=2 w g =1.
[0079] In the embodiment of the present invention, by g Perform Z score standardization to obtain S g N .
[0080] In an embodiment of the present invention, the weight coefficient corresponding to each sliding window length can be obtained based on the following steps:
[0081] When p≤p0, the smaller the length of the sliding window, the larger the weight coefficient. p0 is a preset value. In an exemplary embodiment, p0=5.
[0082] In an exemplary embodiment, when p≤5, if g=2, the corresponding weight coefficient may be 0.4, if g=3, the corresponding weight coefficient may be 0.3, and if g=4 or g=5, the corresponding weight coefficient is 0.15.
[0083] When p>p0, the weight coefficient of the sliding window with a larger length is larger.
[0084] In an illustrative embodiment, when p > 5, the total weight of small windows is 0.4, which is evenly distributed by the number of windows: the weight coefficient corresponding to g = 2 is 0.2 (0.4 ÷ 2), and the weight coefficient corresponding to g = 3 is 0.2 (0.4 ÷ 2). The total weight of large windows is 0.4, which is evenly distributed by the number of windows (fixed at 2): the weight of g = p-1 is 0.2 (0.4 ÷ 2), and the weight of g = p is 0.2 (0.4 ÷ 2). The total weight of medium windows is 0.2, which is evenly distributed by the number of windows (d = p-5, because there are p-5 windows from g = 4 to g = p-2): the weight of each medium window is 0.2 ÷ d. For example, if p = 7, if g = 2 or g = 3, the corresponding weight coefficient can be 0.2, if g = 4 or g = 5, the corresponding weight coefficient can be 0.1, and if g = 6 or g = 7, the corresponding weight coefficient can be 0.2.
[0085] S4102: Input the m sample input data into the n candidate large models respectively to obtain the output result record set of the i-th candidate large model.
[0086] S4103, based on the output result record set of the i-th candidate large model and the corresponding true result, obtain the accuracy record set of the i-th candidate large model.
[0087] In the embodiment of the present invention, the output result of the i-th candidate large model is the predicted answer to the input data. The true result refers to the true answer to the sample input data.
[0088] S4104: Based on the accuracy record set and corresponding feature information of the i-th candidate large model, obtain the question accuracy correlation formula of the i-th candidate large model.
[0089] In the embodiment of the present invention, the problem accuracy correlation formula of the i-th candidate large model can be the multivariate nonlinear regression model A i (p,S)=k0+k1×p+k2×S+k3×p 2 +k4×S 2 +k5×p×S+b. A i (p, S) represents the accuracy of the i-th candidate large model, and S represents the relevance of the problem. k1 to k5 are the parameters to be fitted, and b is the random error term, which satisfies the assumptions of zero mean and constant variance.
[0090] In the embodiment of the present invention, the quadratic term (p 2 , S 2 ) and the interaction term (p×S), which can capture the nonlinear effects of the number of questions (p) and the relevance of questions (S) on the accuracy (such as the marginal attenuation effect of accuracy when the number of questions is too large and the change in the growth rate of accuracy in high-relevance scenarios), which is consistent with the actual change pattern of the accuracy of the large model with the input features.
[0091] In an exemplary embodiment of the present invention, the least square method may be used to solve the parameters. In another embodiment of the present invention, the weighted least square method with L2 regularization may be used to solve the parameters, and the core optimization objective is: .
[0092] Among them, w j is the weight of the j-th sample input data, w j =0.4×S1 j +0.3×S2 j +0.3×S3 j , where S1 j The data volume score of the j-th sample input data ranges from 0 to 5. The larger the data volume, the higher the score. S2 j The annotation quality score of the jth sample input data ranges from 0 to 5. The better the annotation quality, the higher the score. S3 j The domain representativeness score of the jth sample input data ranges from 0 to 5. The more typical it is in the field, the higher the score. j Equal to the number of repeated tests of the j-th sample input data in different scenarios / times. For example, for the sample with "number of questions p=3 and correlation S=0.6", the test was repeated 8 times in different model loading states and different input batches, and the data volume is 8. S2 jIt can be determined based on the consistency of the annotations of the annotators. For example, if there are three annotators, if the annotation results of the three are completely consistent (such as 0.85), 5 points are awarded. If the maximum difference of the annotation results of the three is less than or equal to 0.05 (such as 0.83, 0.85, 0.87), 3 points are awarded. If the maximum difference of the annotation results of the three is greater than 0.1 (such as 0.7, 0.8, 0.9), 1 point is awarded. If there is only one person annotating, 2 points are awarded. If the j-th sample input data is in the core distribution interval of the field to which it belongs, that is, in the interval [Q1 j , Q3 j ], S3 is equal to 5 points, such as in the interval [Q1 j -1.5IQR j , Q3 j +1.5IQR j ], S3 is equal to 3 points, otherwise, S3 is equal to 1 point. j The first quartile of the sample in the field to which the j-th sample input data belongs, Q3 j The third quartile of the sample of the field to which the j-th sample input data belongs, IQR j =Q3 j -Q1 j Among them, Q1 j and Q3 j Obtain it through the following steps:
[0093] (1) Determine the field and sample set
[0094] First, we define the domain label of the j-th sample input data (such as "medical", "legal", "general", etc.), assuming that its domain is D.
[0095] Collect all historical sample data of domain D to form a feature set: XD={x1,…,xa}, where a is the total number of samples in domain D and xa is the core feature value of the ath sample, such as the number of questions or the relevance of questions.
[0096] (2) Data sorting
[0097] Arrange the eigenvalues in the set XD in ascending order to obtain an ordered sequence.
[0098] (3) Calculate quartile positions
[0099] The first quartile (Q1 j ) is calculated as follows: pos1 j =(a+1) / 4, the third quartile (Q3 j ) is calculated as: pos3 j =3(a+1) / 4.
[0100] (4) Quantile value (linear interpolation method)
[0101] If pos1 j or pos3 j is an integer, Q1 j or Q3 j Directly take the value of the corresponding position in the ordered sequence, if pos1 j or pos3 j For decimals, calculate Q1 by linear interpolation j or Q3 j .
[0102] Among them, A ij It represents the actual accuracy of the i-th candidate model on the j-th sample input data, which can be the matching degree between the model output calculated by the text similarity algorithm (such as cosine similarity, edit distance) and the standard answer. i (p j , S j ) is the prediction accuracy of the i-th candidate large model on the j-th sample input data. λ is the regularization coefficient.
[0103] Those skilled in the art should understand that the method of solving the parameters by using the least square method or the weighted least square method with L2 regularization belongs to the scope of the prior art.
[0104] Furthermore, S610 specifically includes:
[0105] S6101, obtain the comprehensive score SC of the i-th candidate large model i =ε×AC i + (1-ε)(1-TC i ).
[0106] Among them, AC i is the normalized value of the accuracy of the i-th candidate large model, AC i =(A i -Amin) / (Amax-Amin), A i is the accuracy of the i-th candidate large model, Amin is the minimum accuracy of all candidate large models, and Amax is the maximum accuracy of all candidate large models. i is the normalized value of the processing time of the i-th candidate large model, TC i =(T i -Tmin) / (Tmax-Tmin), T iis the processing time of the i-th candidate large model, Tmin is the minimum processing time of all candidate large models, and Tmax is the maximum processing time of all candidate large models. ε is the weight coefficient, ranging from 0 to 1, and can be set based on actual conditions. For example, in scenarios where accuracy is prioritized, ε = 0.7 to 0.9, and in scenarios where speed is prioritized, ε = 0.3 to 0.5.
[0107] S6102, the candidate large model with the highest comprehensive score among all candidate large models is used as the target model.
[0108] Those skilled in the art should understand that if the number of candidate large models with the highest comprehensive scores is greater than 1, the candidate large model with the shortest processing time is preferentially selected as the target model.
[0109] In another embodiment of the present invention, S610 specifically includes:
[0110] S6103, if the accuracy of the i-th candidate large model A i < A0, the candidate large model is deleted from the large model record set; the initial value in the large model record set is n candidate large models. A0 is the accuracy threshold, which can be determined based on actual conditions, for example, A0 = 0.85.
[0111] S6104, obtain the comprehensive score SC of the t-th candidate large model in the large model record set t =ε×AC t + (1-ε)(1-TC t ). Among them, AC t is the normalized accuracy value of the t-th candidate large model, where t ranges from 1 to M, and M is the number of models in the large model record set.
[0112] S6105, taking the candidate large model with the highest comprehensive score among all candidate large models in the large model record set as the target model.
[0113] Compared with the above embodiments, this embodiment eliminates large models with accuracy below the threshold, thereby ensuring that the target model is selected more accurately.
[0114] Based on the same inventive concept, an embodiment of the present invention provides a system for optimizing data processing time of a large model, the system comprising:
[0115] The processing time acquisition module is used to input m sample input data into n candidate large models respectively, and obtain the processing time record set T of the i-th candidate large model i ={T i1 , T i2 ,……,T ij ,……,T im}, where T ij It is the processing time of the i-th candidate large model for the j-th sample input data, and the value of j ranges from 1 to m.
[0116] Time character number association acquisition module, used for T i , obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model.
[0117] The discrete degree acquisition module is used to obtain the discrete degree DS corresponding to the jth sample input data based on the processing time record set corresponding to the n candidate large models j The degree of dispersion is used to characterize the degree of fluctuation in the processing time difference when different large models process the same sample input data.
[0118] The target character number acquisition module is used to traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m}, for the discrete degree DS traversed j , if DS j >DS0, change DS j The number of characters of the corresponding sample input data is used as the target number of characters.
[0119] A target model acquisition module is used to obtain the number of characters of the target input data in response to receiving the target input data to be processed. If the number of characters of the target input data is less than the target number of characters, randomly select one large model from the m candidate large models as the target model for processing the target input data; otherwise, input the number of characters of the target input data into the time-character number association formula of each candidate large model to obtain the processing time of each candidate large model for processing the target input data, and use the candidate large model with the minimum processing time as the target model for processing the target input data.
[0120] The system can be used to perform Figure 1 Therefore, for the functions that can be realized by each functional module of the system, please refer to Figure 1 The description of the illustrated embodiment is omitted for brevity.
[0121] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.
[0122] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.
[0123] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.
[0124] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for optimizing data processing time of a large model, characterized in that: The method comprises the following steps: S100: Input m sample input data into n candidate large models respectively to obtain the processing time record set T of the i-th candidate large model. i ={T i1 , T i2 ,……,T ij ,……,T im }, where T ij The processing time of the jth sample input data by the i-th candidate large model, where j ranges from 1 to m; S200, based on T i , obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model; S300, based on the processing time record set corresponding to the n candidate large models, obtain the discrete degree DS corresponding to the j-th sample input data j , the degree of dispersion is used to characterize the degree of fluctuation of the processing time difference when different large models process the same sample input data; S400, traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m }, for the discrete degree DS traversed j , if DS j >DS0, change DS j The number of characters in the corresponding sample input data is used as the target number of characters; DS0 is the threshold value for setting the discrete degree; S500, in response to receiving target input data to be processed, obtaining the number of characters in the target input data; if the number of characters in the target input data is less than the target number of characters, randomly selecting a large model from m candidate large models as the target model for processing the target input data; otherwise, inputting the number of characters in the target input data into a time-to-character-number association formula for each candidate large model to obtain a processing time for each candidate large model to process the target input data; S600: Using the candidate large model with the shortest processing time as the target model for processing the target input data.
2. The method according to claim 1, characterized in that S300 specifically includes: S310, obtaining the processing time record set T corresponding to the j-th sample input data j ={T 1j , T 2j ,……,T ij ,……,T nj }; S320, get T j The corresponding time difference set △T j ={△T j 12 , △T j 23 ,……,△T j (i-1)i ,……,△T j (n-1)n }, △T j (i-1)i is △T j The i-1th time difference in j (i-1)i =T ij -T (i-1)j ; S330, obtain △T j The corresponding degree of dispersion , △T jr is △T j The rth time difference in the time interval, r ranges from 1 to n-1, Avg (△T j ) is △T j The average of the n-1 time differences in .
3. The method according to claim 1, characterized in that After S400, the following steps are also included: S410, based on the m sample input data and the corresponding real results, obtaining a question accuracy correlation formula for each candidate large model, wherein the question accuracy correlation formula is used to represent the relationship between the number of questions in the input data, the question relevance, and the accuracy of the large model; After S500, the following steps are also included: S510, obtaining the number of questions and question relevance in the target input data; S520, inputting the number of questions and the question relevance corresponding to the target input data into the question accuracy correlation formula of each candidate large model to obtain the corresponding accuracy; The S600 is replaced by: S610 , based on the processing time and accuracy of each candidate large model, selecting a target model from the m candidate large models as a target model for processing the target input data.
4. The method according to claim 3, characterized in that S410 specifically includes: S4101, obtaining characteristic information of each sample input data, wherein the characteristic information includes the number of questions and the relevance of the questions; S4102: Input the m sample input data into n candidate large models respectively to obtain the output result record set of the i-th candidate large model; S4103, based on the output result record set of the i-th candidate large model and the corresponding true result, obtain the accuracy record set of the i-th candidate large model; S4104: Based on the accuracy record set and corresponding feature information of the i-th candidate large model, obtain the question accuracy correlation formula of the i-th candidate large model.
5. The method according to claim 4, characterized in that The problem relevance of the sample input data or target input data meets the following conditions: S10, with a sliding window length of g and a sliding step of 1, divide all the problems of the data to be processed into the corresponding problem record set Q g ={Q g1 , Q g2 ,……,Q gh ,……,Q g(z(g)) }, Q gh Q g The h-th problem subset in Q, where h ranges from 1 to z(g), and z(g) is g The number of problem subsets in the dataset; the value of g ranges from 2 to p, where p is the number of problems in the dataset to be processed, and the dataset to be processed is the sample input data or the target input data; S20, get Q gh The corresponding question relevance S gh , and based on the obtained z (g) question relevance, obtain Q g The corresponding question relevance S g ; S30 , obtaining the problem relevance of the data to be processed based on the problem relevance corresponding to the p problem record sets.
6. The method according to claim 5, characterized in that S gh The following conditions must be met: , where sim(Q gh u , Q gh v ) is Q gh The semantic similarity between the u-th question and the v-th question in , where u ranges from h to h+g-1 and v ranges from u+1 to h+g-1.
7. The method according to claim 6, characterized in that S g The following conditions must be met: S g =(∑ z(g) h=1 (w h ×S gh )) / (∑ z(g) h=1 w h ), w h For S gh The weight coefficient of .
8. The method according to claim 6, characterized in that The problem relevance of the data to be processed satisfies the following conditions: S=∑ p g=2 (w g ×S g N ), S g N For S g The normalized value of w g For S g N The corresponding weight coefficient.
9. A system for optimizing data processing time for large models, characterized in that: The system comprises: The processing time acquisition module is used to input m sample input data into n candidate large models respectively, and obtain the processing time record set T of the i-th candidate large model i ={T i1 , T i2 ,……,T ij ,……,T im }, where T ij The processing time of the jth sample input data by the i-th candidate large model, where j ranges from 1 to m; Time character number association acquisition module, used for T i , obtain the time-character-number correlation formula of the i-th candidate large model, which is used to represent the correlation between the processing time and the number of characters of the large model; The discrete degree acquisition module is used to obtain the discrete degree DS corresponding to the jth sample input data based on the processing time record set corresponding to the n candidate large models j , the degree of dispersion is used to characterize the degree of fluctuation of the processing time difference when different large models process the same sample input data; The target character number acquisition module is used to traverse the discrete degree set DS={DS1, DS2, ..., DS j ,……,DS m }, for the discrete degree DS traversed j , if DS j >DS0, change DS j The number of characters in the corresponding sample input data is used as the target number of characters; DS0 is the threshold value for setting the discrete degree; A target model acquisition module is used to obtain the number of characters of the target input data in response to receiving the target input data to be processed. If the number of characters of the target input data is less than the target number of characters, randomly select one large model from the m candidate large models as the target model for processing the target input data; otherwise, input the number of characters of the target input data into the time-character number association formula of each candidate large model to obtain the processing time of each candidate large model for processing the target input data, and use the candidate large model with the minimum processing time as the target model for processing the target input data.
Citation Information
Patent Citations
Business data prediction method and device, electronic equipment and storage medium
CN114202123A
Intelligent optimization method and system based on database performance
CN117290339A