A task-based model inference method and device

Through the task-based model inference method, the index set and task matrix are obtained using task identification, and only the inference layer related to the task is activated, solving the problem that the large model is difficult to take into account inference efficiency and performance in multiple task scenarios, achieving efficient and excellent inference effects.

CN119294530BActive Publication Date: 2025-06-27ZHUO SHI ZHI XING (QINGTIAN) YUAN UNIVERSE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411817302.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-06-27
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The inference efficiency and performance of existing large models in multiple task scenarios are difficult to balance, and model compression is usually aimed at a single task scenario, resulting in performance losses in different task scenarios.

Method used

The task-based model inference method is adopted, by obtaining the task identifier of the to be processed data, obtaining the corresponding index set and task matrix, and using the layer task matrix to transform the data to be processed to obtain the inference result. This method does not require all model layer processing, and only the inference layer related to the task is activated.

Benefits of technology

It realizes supporting efficient inference in multiple task scenarios, taking into account inference efficiency and performance, reducing unnecessary computing burdens and improving inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294530B_ABST
    Figure CN119294530B_ABST
Patent Text Reader

Abstract

The present invention provides a task-based model inference method and apparatus, relating to the field of artificial intelligence technology. The method includes: obtaining data to be processed and a task identifier corresponding to the data to be processed; using the task identifier to obtain an index set corresponding to the data to be processed; using the task identifier to obtain a task matrix corresponding to the data to be processed; and performing transformation processing on the data to be processed with layer task matrices corresponding to all inference layers in an inference model to obtain an inference result of the data to be processed. In the present invention, both the preset index and the layer task matrix are related to the task identifier, which can ensure the inference performance in different task scenarios. Processing the data to be processed with the layer task matrix of the inference layer with a preset index does not require using all model layers for processing, which can effectively reduce unnecessary calculations and improve the inference efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a task-based model inference method and device. Background Art

[0002] With the rapid development of artificial intelligence technology, large models are applied in various task scenarios. A large model is a neural network model with extremely large number of parameters. Due to its huge number of parameters, there is a situation of excessive inference calculation burden during application, which limits the practicability and scalability of the large model, resulting in low inference efficiency when applying the large model.

[0003] Currently, the computational cost of inference can be reduced by means such as model compression. Model compression refers to performing operations such as parameter compression and dimension reduction on the original network structure to improve the inference speed of the network. However, model compression is usually for a single task scenario, and a compressed model will be saved for each task scenario, which is not convenient for maintenance and use. In order to pursue the generality of the model, some model compressions are balanced among multiple task scenarios, resulting in certain performance losses in different task scenarios. Therefore, there is an urgent need for an inference method that can support multiple task scenarios and also take into account inference efficiency and inference performance. Summary of the Invention

[0004] Aiming at the above problems, the purpose of the present invention is to provide a task-based model inference method and device, which can support multiple task scenarios and also take into account inference efficiency and inference performance.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] On the one hand, the present invention provides a task-based model inference method, including:

[0007] Obtain the data to be processed and the task identifier corresponding to the data to be processed;

[0008] Use the task identifier to obtain the index set corresponding to the data to be processed, and the index set includes a plurality of preset indexes;

[0009] Use the task identifier to obtain the task matrix corresponding to the data to be processed, and the task matrix includes layer task matrices associated with the preset indexes;

[0010] Perform transformation processing on the data to be processed with the layer task matrices corresponding to all inference layers in the inference model, and obtain the inference result of the data to be processed, where the layer index of the inference layer in the inference model is the preset index.

[0011] Optionally, transforming the data to be processed by using the layer task matrices corresponding to all the inference layers in the inference model to obtain the inference result of the data to be processed includes:

[0012] For each model layer in the inference model, if the layer input data corresponding to the model layer is received, determine whether the model layer is an inference layer according to the layer index of the model layer in the inference model and the index set, where the layer input data is related to the data to be processed;

[0013] If the model layer is an inference layer, transform the layer input data into layer output data through the layer task matrix corresponding to the inference layer, and transmit the layer output data to the next model layer of the model layer;

[0014] If the model layer is not an inference layer, use the layer input data as the layer output data, and transmit the layer output data to the next model layer of the model layer;

[0015] Use the layer output data of the last model layer in the inference model as the inference result of the data to be processed.

[0016] Optionally, transforming the data to be processed by using the layer task matrices corresponding to all the inference layers in the inference model to obtain the inference result of the data to be processed includes:

[0017] Determine the model layers in the inference model whose layer indices are the same as the preset index as inference layers;

[0018] For each of the inference layers, if the layer input data corresponding to the inference layer is obtained, the layer input data is related to the data to be processed;

[0019] Transform the layer input data into layer output data through the layer task matrix corresponding to the inference layer, and transmit the layer output data to the next inference layer of the inference layer;

[0020] Use the layer output data of the last inference layer in the inference model as the inference result of the data to be processed.

[0021] Optionally, the layer task matrix includes a layer upsampling matrix and a layer downsampling matrix. Transforming the layer input data into layer output data through the layer task matrix corresponding to the inference layer includes:

[0022] Obtain the initial parameter matrix of the inference layer;

[0023] Calculate the product of the layer upsampling matrix and the layer downsampling matrix to obtain an intermediate parameter matrix;

[0024] Sum the initial parameter matrix and the intermediate parameter matrix to obtain an inference parameter matrix;

[0025] Perform a calculation process on the layer input data using the inference parameter matrix to obtain layer output data.

[0026] Optionally, before using the task identifier to obtain the index set corresponding to the data to be processed, the method further includes:

[0027] Obtain a first task data set corresponding to a preset task identifier;

[0028] Perform inference on the first task data set using the inference model to obtain the specified performance parameters of the inference model and the layer input vector of each model layer;

[0029] Calculate the importance parameter of each model layer under the first task data set by using the specified performance parameters and the layer input vector of each model layer;

[0030] Calculate a target number based on the layer input vector of each model layer and the inference resource consumption;

[0031] Determine the target number of inference layers from the inference model according to the importance parameter, and use the layer indexes of all the inference layers as the preset index set corresponding to the preset task identifier.

[0032] Optionally, the performing inference on the first task data set using the inference model to obtain the specified performance parameters of the inference model and the layer input vector of each model layer includes:

[0033] Perform a first inference on the first task data set using all the model layers in the inference model to obtain the benchmark performance parameters of the inference model and the layer input vector of each model layer;

[0034] For each model layer in the inference model, perform a second inference on the first task data set using the other model layers except the model layer to obtain the layer performance parameter corresponding to the model layer;

[0035] Use the benchmark performance parameter and the layer performance parameter as the specified performance parameters of the inference model.

[0036] Optionally, the calculating the importance parameter of each model layer under the first task data set by using the specified performance parameters and the layer input vector of each model layer includes:

[0037] Calculate the relative weight of the model layer by using the initial parameter matrix corresponding to the model layer;

[0038] Calculate the performance change parameter of the model layer based on the reference performance parameter and the layer performance parameter;

[0039] Calculate the feature similarity of the model layer by using the layer input vector of the model layer and the layer input vector of the next model layer;

[0040] Subtract the feature similarity from the sum of the relative weight and the performance change parameter to obtain the importance parameter corresponding to the model layer.

[0041] Optionally, calculating the target quantity based on the layer input vector of each model layer and the inference resource consumption includes:

[0042] Obtain the resource consumption of the first inference and the total deployed resources;

[0043] Calculate the ratio of the resource consumption to the total resources to obtain the resource consumption ratio;

[0044] Calculate the task processing difficulty corresponding to the preset task identifier by using the layer input vector of each model layer;

[0045] Calculate the target quantity according to the resource consumption ratio and the task processing difficulty.

[0046] Optionally, after determining the target number of inference layers from the multiple model layers in the inference model according to the importance parameter and using the indexes of the inference layers as the preset index set corresponding to the preset task identifier, the method further includes:

[0047] For each inference layer, insert an initial dimensionality reduction matrix and an initial dimensionality increase matrix into the initial parameter matrix corresponding to the inference layer;

[0048] Fine-tune the inference model by using the second task data set corresponding to the preset task identifier, and perform gradient update on the initial dimensionality reduction matrix and the initial dimensionality increase matrix to obtain the layer dimensionality reduction matrix and the layer dimensionality increase matrix corresponding to the inference layer;

[0049] Use the layer index corresponding to the inference layer as the preset index, and the layer dimensionality reduction matrix and the layer dimensionality increase matrix as the layer task matrix;

[0050] Use the layer task matrices corresponding to all the preset indexes as the preset task matrix, and establish an association relationship between the preset task matrix and the preset task identifier.

[0051] On the other hand, the present invention further provides a task-based model inference device for implementing the method described in any one of the above, and the device includes:

[0052] A data acquisition module for acquiring the data to be processed and the task identifier corresponding to the data to be processed;

[0053] An index acquisition module, configured to use the task identifier to acquire an index set corresponding to the data to be processed, where the index set includes a plurality of preset indexes;

[0054] A matrix acquisition module, configured to use the task identifier to acquire a task matrix corresponding to the data to be processed, where the task matrix includes layer task matrices associated with the preset indexes;

[0055] An inference module, configured to perform transformation processing on the data to be processed by using layer task matrices corresponding to all inference layers in an inference model, to obtain an inference result of the data to be processed, where a layer index of the inference layer in the inference model is the preset index.

[0056] On the other hand, the present invention further provides an electronic device, including a processor and a memory, where the memory stores a plurality of instructions; the processor loads the instructions from the memory to execute the steps in any one of the task-based model inference methods provided by the present invention.

[0057] On the other hand, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the task-based model inference methods provided by the present invention.

[0058] On the other hand, the present invention further provides a computer program product, including a computer program / instructions, where when the computer program / instructions are executed by a processor, the steps in any one of the task-based model inference methods provided by the present invention are implemented.

[0059] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0060] In an embodiment of the present invention, the data to be processed and a task identifier corresponding to the data to be processed are acquired, and an index set and a task matrix of the data to be processed are acquired by using the task identifier. The index set includes a plurality of preset indexes, and the task matrix includes layer task matrices associated with each preset index. Both the preset indexes and the layer task matrices are related to the task identifier, which can ensure the inference performance in different task scenarios. The layer task matrices of the inference layers with preset indexes are used to process the data to be processed, without using all model layers for processing, which can effectively reduce unnecessary calculations and improve the inference efficiency. Description of the Drawings

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0062] Figure 1 It is a schematic diagram of an application scenario of the task-based model inference method provided by an embodiment of the present invention;

[0063] Figure 2 It is a schematic flowchart of the task-based model inference method provided by an embodiment of the present invention;

[0064] Figure 3 It is a schematic diagram of inferring the first task data set provided by an embodiment of the present invention;

[0065] Figure 4 It is a schematic flowchart of a process for inferring the data to be processed provided by an embodiment of the present invention;

[0066] Figure 5 It is another schematic flowchart of a process for inferring the data to be processed provided by an embodiment of the present invention;

[0067] Figure 6 It is a schematic structural diagram of the task-based model inference device provided by an embodiment of the present invention;

[0068] Figure 7 It is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0070] It can be understood that in the specific implementation manners of the present invention, regarding data such as user information, user permission or consent needs to be obtained, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0071] The present invention proposes a task-based model inference method and device, which can support multiple task scenarios while taking into account inference efficiency and inference performance. For reference, see Figure 1, which shows a schematic diagram of the application scenario of the task-based model inference method. Among them, the application scenario may include a terminal 101 and a server 102. Data exchange can be carried out between the terminal 101 and the server 102 through a network, and a corresponding application program can be installed on the terminal 101. Among them, the terminal 101 can be a mobile phone, a tablet computer, a smart Bluetooth device, a computer, a large screen and other devices, a robot, etc.; the server 102 can be a single server or a server cluster composed of multiple servers.

[0072] The user can send the data to be processed and the corresponding task identifier to the server 102 through the terminal 101. The server 102 can use the task identifier to obtain the index set corresponding to the data to be processed. The index set includes multiple preset indexes; use the task identifier to obtain the task matrix corresponding to the data to be processed. The task matrix includes layer task matrices associated with the preset indexes; use the layer task matrices corresponding to all inference layers in the inference model to perform transformation processing on the data to be processed to obtain the inference result of the data to be processed. Among them, the layer index of the inference layer in the inference model is a preset index.

[0073] The server 102 can send the inference result to the terminal 101 so that the terminal 101 can display the inference result to the user.

[0074] In this embodiment, a task-based model inference method is provided, as Figure 2 shown. The specific process of the task-based model inference method can be as follows:

[0075] S110. Obtain the data to be processed and the task identifier corresponding to the data to be processed.

[0076] The data to be processed is the data that the user inputs and needs to be processed by the model. The data to be processed can be one or more of text, images, and videos. Specifically, the type of the data to be processed can be determined according to the type of data that the model can process.

[0077] The task identifier is the task information associated with the data to be processed and can be input by the user. One task identifier corresponds to one task. For example, a question-and-answer task in the field of health care corresponds to one task identifier, a policy question-and-answer task corresponds to one task identifier, and a question-and-answer task in the legal field corresponds to one task identifier. The corresponding relationship between the task identifier and the task can be pre-agreed.

[0078] S120. Use the task identifier to obtain the index set corresponding to the data to be processed.

[0079] After obtaining the task identifier, the task identifier can be used to obtain the index set corresponding to the data to be processed, where the index set includes multiple preset indexes. Optionally, the first association relationship between the preset task identifier and the preset index set can be obtained first, and then based on the first association relationship, the preset index set corresponding to the task identifier can be obtained as the index set corresponding to the data to be processed.

[0080] Among them, before using the task identifier to obtain the index set corresponding to the data to be processed, the first association relationship between the preset task identifier and the preset index set can be established first. Store the first association relationship between the preset task identifier and the preset index set at a specified location, and it can be directly accessed at the specified location when needed. Based on the first association relationship, the task identifier can be searched in the preset task identifier, and the preset index set corresponding to the task identifier can be used as the index set.

[0081] In some embodiments, when establishing the first association relationship between the preset task identifier and the preset index set, the first task data set corresponding to the preset task identifier can be obtained; the first task data set is inferred by the inference model to obtain the specified performance parameter of the inference model and the layer input vector of each model layer; using the specified performance parameter and the layer input vector of each model layer, calculate the importance parameter of each model layer under the first task data set; based on the layer input vector of each model layer and the inference resource consumption, calculate the target quantity; determine the target number of inference layers from the inference model according to the importance parameter, and use the layer indexes of all the inference layers as the preset index set corresponding to the preset task identifier.

[0082] The preset task identifier can be set according to actual needs, and one preset task identifier corresponds to one task. There can be multiple preset task identifiers. For each preset task identifier, the corresponding first task data set can be obtained, where the first task data set is a small sample calibration data set, usually containing 1024 samples. Using the inference model to infer the first task data set, the performance parameter of the inference model and the layer input vector of each model layer in the inference model can be obtained.

[0083] The inference model is a large model, which can include an embedding layer and multiple attention layers. The model layer referred to in the embodiments of the present invention is the attention layer. The specified performance parameter can be accuracy, perplexity, etc., which can be specifically set according to actual needs. When inferring the first task data set, it can be inferred multiple times, and the model layers used in each inference are different to evaluate the influence of each model layer on the inference. In the embodiments of the present invention, the number of inferences is related to the total number of model layers in the inference model.

[0084] When using the inference model to perform inference on the first task dataset, each data in the first task dataset will pass through the model layers, so that the layer input vectors corresponding to each model layer can be obtained. Using the specified performance parameters and the layer input vectors of each model layer, the importance parameters of each model layer can be calculated under this first task dataset. The importance parameter can reflect the importance degree of each model layer in processing the first task dataset. The larger the importance parameter, the more important the model layer is. Under different tasks, the participation degrees of each model layer in the large model in the inference are different. Evaluating the importance parameters of each model layer based on the first task dataset is more suitable for the specific task referred to by the preset task identifier.

[0085] The target number refers to the number of model layers required to be used when inferring the task referred to by the preset task identifier. During the model inference process, it is affected by computing resources and task difficulty. Therefore, the target number can be comprehensively determined based on the layer input vectors of each model layer and the inference resource consumption. Then, select the target number of inference layers from the inference model according to the importance parameters, and use the indices of all the inference layers as the preset index set corresponding to the preset task identifier.

[0086] In order to comprehensively evaluate the contribution degree of each model layer to the inference, when using the inference model to perform inference on the first task dataset, it can be to use all the model layers in the inference model to perform the first inference on the first task dataset to obtain the baseline performance parameter of the inference model and the layer input vectors of each model layer; for each model layer in the inference model, use the other model layers except this model layer to perform the second inference on the first task dataset to obtain the layer performance parameter corresponding to this model layer; use the baseline performance parameter and the layer performance parameter as the specified performance parameters of the inference model.

[0087] For example, refer to Figure 3 which shows a schematic diagram of inferring the first task dataset. Among them, the first inference means using all the model layers in the inference model to perform inference on the first task dataset. The task dataset can be sequentially processed by multiple model layers to obtain the final inference result. During this process, the layer input vectors corresponding to each model layer and the performance parameter of the inference model can be obtained, and this performance parameter can be used as the baseline performance parameter.

[0088] The second inference means removing one of the model layers in the inference model and using the remaining model layers to perform inference on the task dataset. Among them, removing can refer to methods such as masking, skipping, and eliminating. For example, Figure 3Among them, the dashed box represents the removed model layer, and the data directly enters the subsequent model layer without being processed by this model layer. If there are n model layers in the inference model, n second inferences need to be performed. One model layer is removed during each second inference. During each second inference, the performance parameters of the inference model can be recorded. Since one model layer is removed during each second inference, the performance parameter can be recorded as the layer performance parameter of the model layer. After n second inferences, the layer performance parameters of each model layer can be obtained. Finally, the baseline performance parameter and the layer performance parameters of each model layer can be used as the specified performance parameters of the inference model.

[0089] The specified performance parameters can include the baseline performance parameter and multiple layer performance parameters, and the layer performance parameters correspond one-to-one with the model layers in the inference model. Using the specified performance parameters and the layer input vectors of each model layer, the importance parameters of each model layer can be calculated under the first task dataset. When calculating the importance parameter of a model layer, the relative weight of the model layer can be calculated using the initial parameter matrix corresponding to the model layer; the performance change parameter of the model layer can be calculated using the baseline performance parameter and the layer performance parameter; the feature similarity of the model layer can be calculated using the layer input vector of the model layer and the layer input vector of the next model layer; the importance parameter corresponding to the model layer can be obtained by subtracting the feature similarity from the sum of the relative weight and the performance change parameter.

[0090] Each model layer in the inference model has its corresponding initial parameter matrix, which can also be called the weight matrix. This initial parameter matrix is obtained through a large amount of training. The model weight accuracy range defines the value range of the weight values in the model. The model weight accuracy range can include the minimum accuracy value and the maximum accuracy value. Using the minimum accuracy value, the maximum accuracy value, and the initial parameter matrix of this model layer, the relative weight of this model layer can be calculated.

[0091] The aforementioned baseline performance parameter is the performance parameter obtained when all model layers are used for inference, while the layer performance parameter is the performance parameter obtained when inference is performed with the corresponding model layer missing. The degree of influence on the performance of the inference model after the corresponding model layer is missing can be calculated to obtain the performance change parameter.

[0092] Among them, the performance change parameter can first subtract the layer performance parameter from the benchmark performance parameter to obtain the specific value of the performance change, and then divide the specific value of the performance change by the benchmark performance parameter to obtain the performance degradation ratio, which is used as the performance change parameter. Since there are positive and negative indicators for performance parameters, for example, accuracy is a positive indicator, that is, the greater the accuracy, the better the performance, while perplexity is a negative indicator, and the greater the perplexity, the worse the performance. Therefore, a sign parameter can be introduced. When the performance parameter uses a positive indicator, the sign parameter is positive, and when the performance parameter uses a negative indicator, the sign parameter is negative.

[0093] In the first inference, the layer input vector corresponding to each model layer can be obtained, and the average cosine similarity can be calculated based on the layer input vector of this model layer and the layer input vector of the next model layer as the feature similarity of the model layer. The higher the cosine similarity, the higher the feature similarity, indicating that the data content processed by this model layer is less, then the gap between the input vector of this model layer and the input vector of the next model layer is smaller, and the importance of this model layer is lower.

[0094] Then, the relative weight and the performance change parameter can be summed, and then the feature similarity can be subtracted from the sum of the two to obtain the importance parameter corresponding to the model layer.

[0095] Specifically, the importance parameter of the model layer can be calculated with reference to the following formula:

[0096] ;

[0097] Among them, LS j represents the importance parameter of the j-th model layer; represents the L2 norm of the initial parameter matrix of the j-th model layer; w min represents the minimum value of precision; w max represents the maximum value of precision; represents the difference between the benchmark performance parameter and the layer performance parameter of the j-th model layer; metric0 represents the benchmark performance parameter; β is the sign parameter; x j represents the layer input vector of the j-th model layer; x j+1 represents the layer input vector of the (j + 1)-th model layer. Among them, the first term is the relative weight. During the calculation of the relative weight, logarithmic mapping is performed to prevent the data from being too skewed, and linear normalization can map this relative weight to the range of [0, 1], which is consistent with the value ranges of the performance change parameter and the feature similarity.

[0098] Thus, the importance parameter of each model layer can be calculated, and the model layers can be sorted in descending order according to the importance parameter to obtain the model layer sequence.

[0099] In different tasks, the contribution of each model layer is different. To reduce unnecessary computational burden, during the inference task, some model layers can be adaptively activated for inference. As an implementation, an importance threshold can be set, and the model layers with importance parameters greater than the importance threshold are used as inference layers. The inference layers are the model layers that need to be activated when processing data.

[0100] As another implementation, the number of layers to be activated during inference can be comprehensively determined by combining the actual inference resources and task difficulty. For example, it can be to obtain the resource consumption of the first inference and the total deployed resources; calculate the ratio of the resource consumption to the total resources to obtain the resource consumption ratio; calculate the task processing difficulty corresponding to the preset task identifier based on the layer input vector of each model layer; and calculate the target number according to the resource consumption ratio and the task processing difficulty.

[0101] In the first inference and multiple second inferences of the inference model, since different model layers are used, the resources consumed are also different. In the first inference, all model layers are used for inference, and the resource consumption of the first inference can be obtained for subsequent calculations. The resources in the embodiments of the present invention can be understood as computing resources. The total deployed resources refer to the total computing resources on the device for inference. Dividing the resource consumption by the total resources can obtain the resource consumption ratio.

[0102] Different task identifiers correspond to different tasks. Different tasks have different difficulties and different consumptions of computing resources, and the number of model layers used is also different. In the first inference, the layer input vector of each model layer can be obtained. The more similar the layer input vectors between every two adjacent model layers are, the smaller the task difficulty. Considering the consumption of computing resources and the task processing difficulty comprehensively, the target number can be calculated based on the following formula:

[0103] ;

[0104] where k i represents the target number; R represents the total deployed resources on the device; R0 represents the resource consumption of the first inference; n represents the total number of model layers in the inference model; x j represents the layer input vector of the j-th model layer; x j+1 represents the layer input vector of the (j + 1)-th model layer.

[0105] Then, the top k i model layers can be selected from the model layer sequence, and these k i model layers are used as inference layers. Then, the layer index corresponding to each inference layer in the inference model is obtained, and this layer index is stored as a preset index set, and this preset index set is associated and stored with the preset task identifier.

[0106] S130. Obtain the task matrix corresponding to the data to be processed by using the task identifier.

[0107] Based on the task identifier of the data to be processed, the task matrix corresponding to the task identifier can be obtained. The task matrix contains the layer task matrices associated with the preset index. Similarly, a second association relationship between the preset task identifier and the preset task matrix can be established in advance, and each preset task matrix contains the layer task matrices associated with the preset index. Based on this second association relationship, use the task identifier to find the corresponding preset task matrix and use it as the task matrix of the data to be processed.

[0108] Optionally, after establishing the first association relationship between the preset task identifier and the preset index set, in order to further improve the performance of the inference model under specific tasks, the inference layer can be fine-tuned using the second task dataset, and the relevant parameters can be saved.

[0109] Based on the first association relationship established above between the preset task identifier and the preset index set, for a specific task identifier, the layer index of the inference layer can already be obtained. During fine-tuning, for each inference layer, an initial dimensionality reduction matrix and an initial dimensionality increase matrix can be inserted into the initial parameter matrix corresponding to the inference layer; the inference model is fine-tuned using the second task dataset corresponding to the preset task identifier, and the initial dimensionality reduction matrix and the initial dimensionality increase matrix are subjected to gradient update to obtain the layer dimensionality reduction matrix and the layer dimensionality increase matrix corresponding to the inference layer; the layer index corresponding to the inference layer is used as the preset index, and the layer dimensionality reduction matrix and the layer dimensionality increase matrix are used as the layer task matrix; the layer task matrices corresponding to all preset indexes are used as the preset task matrix, and the association relationship between the preset task matrix and the preset task identifier is established.

[0110] The initial parameter matrix is the parameter matrix of each model layer in the inference model before fine-tuning. The initial dimensionality reduction matrix and the initial dimensionality increase matrix are a set of learnable parameter matrices. If the initial parameter matrix is dimensional, then the initial dimensionality reduction matrix is dimensional, and the initial dimensionality increase matrix is dimensional, where r is much smaller than the minimum of d and k. Among them, the initial dimensionality reduction matrix is initialized using a random Gaussian distribution, and the initial dimensionality increase matrix is initially a zero matrix. After inserting the initial dimensionality reduction matrix and the initial dimensionality increase matrix, the forward propagation process can be changed from y = W0x to . Where x is a d-dimensional input vector, y is a k-dimensional output vector, W0 is the initial parameter matrix, A is the initial dimensionality reduction matrix, and B is the initial dimensionality increase matrix.

[0111] The second task dataset corresponding to the preset task identifier refers to the training dataset in this task scenario. The difference between it and the first task dataset is that the data volume of the second task dataset is larger. During the fine-tuning process, each initial parameter matrix in the inference model needs to be frozen, that is, set to non-differentiable. Then, use the second task dataset to train the inference model, calculate the gradients of each group of initial dimensionality reduction matrices and initial dimensionality increase matrices, and update the initial dimensionality reduction matrices and initial dimensionality increase matrices according to the gradients until the convergence condition is met. When the training is completed, the layer dimensionality reduction matrix and layer dimensionality increase matrix corresponding to each inference layer can be obtained.

[0112] The layer dimensionality reduction matrix and layer dimensionality increase matrix can be denoted as the layer task matrix. This layer task matrix corresponds to the inference layer. The layer index of the inference layer can be used as the preset index to establish the association relationship, that is, the second association relationship, between the preset index and the layer task matrix. Since there are multiple inference layers, the association relationships between all preset indexes and layer task matrices can be used as the preset task matrix, and the second association relationship between the preset task matrix and the preset task identifier is established and stored.

[0113] Therefore, after obtaining the task identifier of the data to be processed, the preset task matrix corresponding to the task identifier can be found as the task matrix of the data to be processed. This task matrix contains the layer task matrices corresponding to each preset index.

[0114] S140: Use the layer task matrices corresponding to all inference layers in the inference model to perform transformation processing on the data to be processed to obtain the inference result of the data to be processed.

[0115] An inference layer refers to the model layer in the inference model used to process the data to be processed. The layer index of the inference layer in the inference model is in the preset index set. That is, in the inference model, the model layer with this preset index is the inference layer. Since the preset index is associated with the layer task matrix, the layer task matrix of the corresponding inference layer can be obtained.

[0116] The data to be processed is input into the inference model. After being processed by all model layers in the inference model, the inference result can be obtained. That is, the layer task matrices corresponding to all inference layers can be used to perform transformation processing on the data to be processed to obtain the inference result.

[0117] As an implementation manner, when using the layer task matrices corresponding to all the inference layers in the inference model to perform transformation processing on the data to be processed to obtain the inference result of the data to be processed, for each model layer in the inference model, if the layer input data corresponding to the model layer is received, it is determined whether the model layer is an inference layer according to the layer index of the model layer in the inference model and the index set, and the layer input data is related to the data to be processed; if the model layer is an inference layer, the layer input data is transformed into layer output data through the layer task matrix corresponding to the inference layer, and the layer output data is passed to the next model layer of the model layer; if the model layer is not an inference layer, the layer input data is used as the layer output data, and the layer output data is passed to the next model layer of the model layer; the layer output data of the last model layer in the inference model is used as the inference result of the data to be processed.

[0118] The data to be processed is input into the inference model, and the data to be processed can sequentially pass through each model layer in each inference model to realize inferring the data to be processed by using the inference layer. Reference can be made to Figure 4 , which shows a schematic flowchart of inferring the data to be processed. Among them, for each model layer, if the model layer receives the corresponding layer input data, it can be determined what operation to perform on the layer input data based on the layer index of the model layer and the index set.

[0119] The layer input data refers to the data received by the model layer, and the layer input data is obtained by transforming the data to be processed. When the model layer receives the layer input data, it can determine whether the layer index of the model layer in the inference model is in the index set to determine whether the model layer is an inference layer.

[0120] If the layer index of the model layer in the inference model is in the index set, it can be determined that the model layer is an inference layer; if the layer index of the model layer in the inference model is not in the index set, it can be determined that the model layer is not an inference layer.

[0121] If the model layer is an inference layer, then the model layer needs to process the data. The layer input data can be transformed into layer output data through the layer task matrix corresponding to the inference layer. Then, the layer output data can be passed to the next model layer as the layer input data of the next model layer. If the model layer is not an inference layer, then the model layer does not need to process the data, and the layer input data can be directly used as the layer output data and passed to the next model layer as the layer input data of the next model layer. In this case, the layer input data of the next model layer is the same as the layer input data of this model layer. For example, Figure 4Among them, Data 1 is the input data of the first layer of the model layer; if 1 is in the index set, after being processed by the first layer of the model layer, Data 2 can be obtained as the input data of the second layer of the model layer; if 1 is not in the index set, Data 1 is directly passed to the second layer of the model layer, then Data 1 and Data 2 are the same.

[0122] In this way, by traversing each model layer, the output data of the last model layer in the inference model can be obtained, and this output data is the inference result of the data to be processed.

[0123] A task selection mechanism can be introduced before the embedding layer in the inference model, and the core components include this first association relationship and the message sending program. When a task identifier is received, the index set can be obtained based on the first association relationship and sent to each model layer. A routing mechanism can be introduced before each model layer, and the core components include the message receiving program and the gating program. The message receiving program is used to receive the index set, and the gating program is to judge whether the index of this model layer is in the index set; if it is, route the layer input data to this model layer; if not, route the layer input data to the next model layer.

[0124] As another implementation manner, when using the layer task matrices corresponding to all the inference layers in the inference model to perform transformation processing on the data to be processed to obtain the inference result of the data to be processed, it can be to determine the model layers in the inference model whose layer indices are consistent with the preset index as the inference layers; for each of the inference layers, obtain the layer input data corresponding to the inference layer, and the layer input data is related to the data to be processed; transform the layer input data into layer output data through the layer task matrix corresponding to the inference layer and pass the layer output data to the next inference layer of the inference layer; use the layer output data of the last inference layer in the inference model as the inference result of the data to be processed.

[0125] All the inference layers can be determined from multiple model layers in the inference model according to the index set. For example, refer to Figure 5 , which shows another schematic flowchart of inferring the data to be processed. Among them, the model layers with layer indices 1, 3, 4, n are all inference layers. After the data 1 is processed by the inference layer 1, the layer output data is directly passed to the inference layer 3 for processing without being passed to the inference layer 2. Specifically, the layer indices of each model layer in the inference model can be obtained first, and the model layers whose layer indices are consistent with the preset index in the index set are determined as the inference layers.

[0126] For each inference layer, the layer input data corresponding to the inference layer can be obtained, and the layer input data can be transformed from the data to be processed. Using the layer task matrix of the inference layer, the layer input data of the inference layer can be transformed into layer output data, and the layer output data is passed to the next inference layer of the inference layer, that is, the layer input data of the next inference layer is the layer output data of the inference layer. Finally, the layer output data of the last inference layer in the inference model can be obtained as the inference result of the data to be processed. For example Figure 5 In Figure 5 , assuming that the preset indexes are 1, 3, 4, n, and the data 1 is the layer input data of the model layer 1.

[0127] When transforming the layer input data into layer output data by using the layer task matrix of the inference layer, it may be to obtain the initial parameter matrix of the inference layer; calculate the product of the layer up-dimension matrix and the layer down-dimension matrix to obtain an intermediate parameter matrix; sum the initial parameter matrix and the intermediate parameter matrix to obtain an inference parameter matrix; perform calculation processing on the layer input data with the inference parameter matrix to obtain the layer output data.

[0128] Among them, the layer task matrix contains a layer up-dimension matrix and a layer down-dimension matrix. After the inference model is trained, each model layer has its corresponding initial parameter matrix. Thus, the initial parameter matrix of the inference layer can be obtained. The product of multiplying the layer up-dimension matrix and the layer down-dimension matrix is used as the intermediate parameter matrix, and then the intermediate parameter matrix and the initial parameter matrix are added together to obtain the inference parameter matrix. Finally, calculation processing is performed on the layer input data with the inference parameter matrix to obtain the layer output data.

[0129] When the task identifier does not change, but there is new data to be processed input, the above S140 can be repeated to continue the inference. If the task identifier changes, it is necessary to re-obtain the index set corresponding to the new task identifier to continue the subsequent process. Among them, since the task matrix corresponding to the previous task identifier has been loaded, it is necessary to unload the task matrix first and then load the task matrix corresponding to the new task identifier. For each inference layer, the process of loading the task matrix is to add the product of the layer down-dimension matrix and the layer up-dimension matrix to the basis of the initial parameter matrix, and the unloading process is to subtract the product of the layer down-dimension matrix and the layer up-dimension matrix on the basis of the loading.

[0130] In the embodiments of the present invention, both the model layer and the inference layer are transformer layers. In the transformer architecture, the sub-attention layer captures the dependencies between elements by calculating the attention scores of each element in the input data of the calculation layer, which involves a large number of matrix multiplication operations. In the text generation task, the model generates each token of the text one by one. During this process, the key and value vectors of the processed tokens can be stored, i.e., the KV cache, for repeated use when generating subsequent tokens, avoiding repeated calculations. In the embodiments of the present invention, each token is processed using the same inference layer, which can ensure the effectiveness of the KV cache, reduce repeated calculations, and improve the inference efficiency.

[0131] The task-based model inference solution provided by the embodiments of the present invention can be applied to various inference scenarios using large models. For example, taking health Q&A as an example, adopting the solution provided by the embodiments of the present invention can adapt to the health Q&A task, only activate the corresponding model layer from the inference model, without using all model layers, which can effectively reduce the inference cost and improve the inference efficiency.

[0132] Through the method provided by the embodiments of the present invention, the task data set can be inferred by the inference model in advance to calculate the importance parameters of each model layer of the inference model layer. Then, based on the inference calculation cost and task difficulty, the number of model layers to be activated is determined, and the model layer used for inference is determined using the importance parameters to obtain the inference layer. In actual inference applications, the corresponding inference layer can be selected based on the task identifier to infer the data, which can effectively improve the inference efficiency. Since the inference layer is determined based on the importance parameters, the inference performance can also be ensured, thus achieving both adaptation to a variety of different task scenarios and consideration of inference performance while improving the inference efficiency.

[0133] To better implement the above method, the embodiments of the present invention also provide a task-based model inference device. The task-based model inference device can be specifically integrated in an electronic device, and the electronic device can be a terminal, a server, or other devices. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or other devices; the server can be a single server or a server cluster composed of multiple servers.

[0134] For example, in this embodiment, taking the task-based model inference device specifically integrated in the server as an example, the method of the embodiments of the present invention will be described in detail.

[0135] For example, as Figure 6 shown, the task-based model inference device 200 may include a data acquisition module 210, an index acquisition module 220, a matrix acquisition module 230, and an inference module 240.

[0136] A data acquisition module 210, configured to acquire data to be processed and a task identifier corresponding to the data to be processed;

[0137] An index acquisition module 220, configured to use the task identifier to acquire an index set corresponding to the data to be processed, where the index set includes a plurality of preset indexes;

[0138] A matrix acquisition module 230, configured to use the task identifier to acquire a task matrix corresponding to the data to be processed, where the task matrix includes layer task matrices associated with the preset indexes;

[0139] An inference module 240, configured to perform transformation processing on the data to be processed with layer task matrices corresponding to all inference layers in an inference model to obtain an inference result of the data to be processed, where a layer index of the inference layer in the inference model is the preset index.

[0140] In some embodiments, the inference module 240 is specifically configured to:

[0141] For each model layer in the inference model, if layer input data corresponding to the model layer is received, determine whether the model layer is an inference layer according to the layer index of the model layer in the inference model and the index set, where the layer input data is related to the data to be processed;

[0142] If the model layer is an inference layer, transform the layer input data into layer output data through the layer task matrix corresponding to the inference layer, and transmit the layer output data to the next model layer of the model layer;

[0143] If the model layer is not an inference layer, use the layer input data as layer output data, and transmit the layer output data to the next model layer of the model layer;

[0144] Use the layer output data of the last model layer in the inference model as the inference result of the data to be processed.

[0145] In some embodiments, the inference module 240 is specifically configured to:

[0146] Determine a model layer in the inference model whose layer index is consistent with the preset index as an inference layer;

[0147] For each of the inference layers, if layer input data corresponding to the inference layer is acquired, where the layer input data is related to the data to be processed;

[0148] Transform the layer input data into layer output data through the layer task matrix corresponding to the inference layer, and transmit the layer output data to the next inference layer of the inference layer;

[0149] Use the layer output data of the last inference layer in the inference model as the inference result of the data to be processed.

[0150] In some embodiments, the layer task matrix includes a layer upsampling matrix and a layer downsampling matrix. The inference module 240 is specifically configured to:

[0151] Obtain the initial parameter matrix of the inference layer;

[0152] Calculate the product of the layer upsampling matrix and the layer downsampling matrix to obtain an intermediate parameter matrix;

[0153] Sum the initial parameter matrix and the intermediate parameter matrix to obtain an inference parameter matrix;

[0154] Perform a calculation process on the layer input data using the inference parameter matrix to obtain layer output data.

[0155] In some embodiments, the task-based model inference device 200 further includes a preparation module. Before obtaining the index set corresponding to the data to be processed by using the task identifier, the preparation module is specifically configured to:

[0156] Obtain a first task data set corresponding to a preset task identifier;

[0157] Perform inference on the first task data set using the inference model to obtain the specified performance parameters of the inference model and the layer input vectors of each model layer;

[0158] Calculate the importance parameters of each model layer under the first task data set by using the specified performance parameters and the layer input vectors of each model layer;

[0159] Calculate a target quantity based on the layer input vectors of each model layer and the inference resource consumption;

[0160] Determine the target quantity of inference layers from the inference model according to the importance parameters, and use the layer indices of all the inference layers as the preset index set corresponding to the preset task identifier.

[0161] In some embodiments, the preparation module is specifically configured to:

[0162] Perform a first inference on the first task data set by using all the model layers in the inference model to obtain the benchmark performance parameters of the inference model and the layer input vectors of each model layer;

[0163] For each model layer in the inference model, perform a second inference on the first task data set by using the other model layers except the model layer to obtain the layer performance parameters corresponding to the model layer;

[0164] Use the benchmark performance parameter and the layer performance parameter as the specified performance parameters of the inference model.

[0165] In some embodiments, the preparation module is specifically configured to:

[0166] Calculate the relative weight of the model layer using the initial parameter matrix corresponding to the model layer;

[0167] Calculate the performance change parameter of the model layer using the benchmark performance parameter and the layer performance parameter;

[0168] Calculate the feature similarity of the model layer using the layer input vector of the model layer and the layer input vector of the next model layer;

[0169] Obtain the importance parameter corresponding to the model layer by subtracting the feature similarity from the sum of the relative weight and the performance change parameter.

[0170] In some embodiments, the preparation module is specifically configured to:

[0171] Obtain the resource consumption of the first inference and the total deployed resources;

[0172] Calculate the ratio of the resource consumption to the total resources to obtain the resource consumption ratio;

[0173] Calculate the task processing difficulty corresponding to the preset task identifier using the layer input vector of each model layer;

[0174] Calculate the target quantity according to the resource consumption ratio and the task processing difficulty.

[0175] In some embodiments, after determining the target number of inference layers from the multiple model layers in the inference model according to the importance parameter and using the indices of the inference layers as the preset index set corresponding to the preset task identifier, the preparation module is specifically configured to:

[0176] For each inference layer, insert an initial dimensionality reduction matrix and an initial dimensionality increase matrix into the initial parameter matrix corresponding to the inference layer;

[0177] Fine-tune the inference model using the second task dataset corresponding to the preset task identifier, and perform gradient update on the initial dimensionality reduction matrix and the initial dimensionality increase matrix to obtain the layer dimensionality reduction matrix and the layer dimensionality increase matrix corresponding to the inference layer;

[0178] Use the layer index corresponding to the inference layer as the preset index, and the layer dimensionality reduction matrix and the layer dimensionality increase matrix as the layer task matrix;

[0179] Take the layer task matrices corresponding to all preset indexes as the preset task matrix, and establish an association relationship between the preset task matrix and the preset task identifier.

[0180] In specific implementation, each of the above modules can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0181] As can be seen from the above, the task-based model inference device in this embodiment can obtain the data to be processed and the task identifier corresponding to the data to be processed, use the task identifier to obtain the index set and task matrix of the data to be processed. The index set contains multiple preset indexes, and the task matrix contains the layer task matrices associated with each preset index. Both the preset index and the layer task matrix are related to the task identifier, which can ensure the inference performance in different task scenarios. The layer task matrix of the inference layer with the preset index is used to process the data to be processed, without using all model layers for processing, which can effectively reduce unnecessary calculations and improve the inference efficiency.

[0182] An embodiment of the present invention further provides an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0183] In some embodiments, the task-based model inference device can also be integrated in multiple electronic devices. For example, the task-based model inference device can be integrated in multiple servers, and the task-based model inference method of the present invention is implemented by multiple servers.

[0184] In this embodiment, the electronic device in this embodiment is taken as an example of a server for detailed description. For example, as Figure 7 shown, it shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:

[0185] The electronic device may include a processor 310 with one or more processing cores, a memory 320 with one or more computer-readable storage media, a power supply 330, an input module 340, and a communication module 350 and other components. Those skilled in the art can understand that Figure 7 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine some components, or different component arrangements. Among them:

[0186] The processor 310 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 320, and by invoking the data stored in the memory 320, it performs various functions of the electronic device and processes data. In some embodiments, the processor 310 may include one or more processing cores; in some embodiments, the processor 310 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 310 either.

[0187] The memory 320 can be used to store software programs and modules. The processor 310 executes various functional applications and data processing by running the software programs and modules stored in the memory 320. The memory 320 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 320 may also include a memory controller to provide the processor 310 with access to the memory 320.

[0188] The electronic device also includes a power supply 330 that powers each component. In some embodiments, the power supply 330 can be logically connected to the processor 310 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 330 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0189] The electronic device may also include an input module 340, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0190] The electronic device may also include a communication module 350. In some embodiments, the communication module 350 may include a wireless module. The electronic device can perform short-distance wireless transmission through the wireless module of the communication module 350, thereby providing users with wireless broadband Internet access. For example, the communication module 350 can be used to help users send and receive emails, browse the web, and access streaming media, etc.

[0191] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 310 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 320 according to the following instructions, and the processor 310 will run the application programs stored in the memory 320, so as to implement the steps in the methods of the embodiments of the present invention.

[0192] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0193] As can be seen from the above, the electronic device provided by the embodiments of the present invention can obtain the data to be processed and the task identifier corresponding to the data to be processed, and use the task identifier to obtain the index set and task matrix of the data to be processed. The index set contains multiple preset indexes, and the task matrix contains the layer task matrices associated with each preset index. Both the preset index and the layer task matrix are related to the task identifier, which can ensure the inference performance in different task scenarios. The layer task matrix of the inference layer with the preset index is used to process the data to be processed, without using all model layers for processing, which can effectively reduce unnecessary calculations and improve the inference efficiency.

[0194] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0195] Therefore, an embodiment of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the task-based model inference methods provided by the embodiments of the present invention.

[0196] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0197] According to an aspect of the present invention, there is provided a computer program product or computer program, which includes computer programs / instructions, and the computer programs / instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer programs / instructions from the computer-readable storage medium, and the processor executes the computer programs / instructions, so that the electronic device executes the methods provided in the various optional implementation manners of the preparation aspect or the application inference aspect provided in the above embodiments.

[0198] Since the instructions stored in the storage medium can execute the steps in any of the task-based model inference methods provided by the embodiments of the present invention, the beneficial effects achievable by any of the task-based model inference methods provided by the embodiments of the present invention can be realized. For details, see the previous embodiments and will not be elaborated here.

[0199] The above has introduced in detail a task-based model inference method and apparatus provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A task-based model reasoning method, characterized in that: The method comprises: Acquire data to be processed and a task identifier corresponding to the data to be processed, wherein the data to be processed is one or more of text, image, and video; Using a first association relationship between a preset task identifier and a preset index set and the task identifier, an index set corresponding to the data to be processed is acquired, wherein the index set includes a plurality of preset indexes; Using the task identifier, obtaining a task matrix corresponding to the data to be processed, wherein the task matrix includes a layer task matrix associated with the preset index; The data to be processed is transformed using the layer task matrices corresponding to all the inference layers in the inference model to obtain the inference result of the data to be processed, and the layer index of the inference layer in the inference model is the preset index; The first association relationship is obtained through the following steps: Obtain a first task data set corresponding to a preset task identifier; perform a first reasoning on the first task data set using all model layers in the reasoning model to obtain a baseline performance parameter of the reasoning model and a layer input vector of each model layer; for each model layer in the reasoning model, perform a second reasoning on the first task data set using other model layers except the model layer to obtain a layer performance parameter corresponding to the model layer; use the baseline performance parameter and the layer performance parameter as the specified performance parameters of the reasoning model; use the specified performance parameter and the layer input vector of each model layer to calculate the importance parameter of each model layer under the first task data set; calculate a target number based on the layer input vector of each model layer and the reasoning resource consumption; determine a target number of reasoning layers from multiple model layers in the reasoning model according to the importance parameter, and use the index of the reasoning layer as the preset index set corresponding to the preset task identifier.

2. The method according to claim 1, characterized in that The transforming process of the data to be processed using the layer task matrix corresponding to all the inference layers in the inference model to obtain the inference result of the data to be processed includes: For each model layer in the inference model, if layer input data corresponding to the model layer is received, determine whether the model layer is an inference layer according to the layer index of the model layer in the inference model and the index set, and the layer input data is related to the data to be processed; If the model layer is an inference layer, the layer input data is transformed into layer output data through a layer task matrix corresponding to the inference layer, and the layer output data is passed to the next model layer of the model layer; If the model layer is not an inference layer, using the layer input data as layer output data, and passing the layer output data to the next model layer of the model layer; The layer output data of the last model layer in the inference model is used as the inference result of the data to be processed.

3. The method according to claim 1, characterized in that The transforming process of the data to be processed using the layer task matrix corresponding to all the inference layers in the inference model to obtain the inference result of the data to be processed includes: Determine, in the inference model, a model layer whose layer index is consistent with the preset index as the inference layer; For each of the inference layers, if layer input data corresponding to the inference layer is obtained, the layer input data is related to the data to be processed; Transforming the layer input data into layer output data through the layer task matrix corresponding to the reasoning layer, and passing the layer output data to the next reasoning layer of the reasoning layer; The layer output data of the last inference layer in the inference model is used as the inference result of the data to be processed.

4. The method according to claim 2 or 3, characterized in that: The layer task matrix includes a layer dimension increase matrix and a layer dimension reduction matrix, and the layer input data is transformed into layer output data by using the layer task matrix corresponding to the reasoning layer, including: Obtaining an initial parameter matrix of the inference layer; Calculate the product of the layer dimension increase matrix and the layer dimension reduction matrix to obtain an intermediate parameter matrix; Summing the initial parameter matrix and the intermediate parameter matrix to obtain an inference parameter matrix; The layer input data is calculated and processed using the inference parameter matrix to obtain layer output data.

5. The method according to claim 1, characterized in that The step of calculating the importance parameter of each model layer under the first task data set by using the specified performance parameter and the layer input vector of each model layer includes: Calculating the relative weight of the model layer using the initial parameter matrix corresponding to the model layer; Calculating a performance change parameter of the model layer using the baseline performance parameter and the layer performance parameter; Calculating the feature similarity of the model layer using the layer input vector of the model layer and the layer input vector of the next model layer; The importance parameter corresponding to the model layer is obtained by subtracting the feature similarity from the sum of the relative weight and the performance change parameter.

6. The method according to claim 1, characterized in that The target quantity is calculated based on the layer input vector of each model layer and the inference resource consumption, including: Obtaining resource consumption of the first inference and total resources deployed; Calculating the ratio of the resource consumption to the total resource amount to obtain a resource consumption ratio; Calculating the task processing difficulty corresponding to the preset task identifier using the layer input vector of each model layer; The target quantity is calculated according to the resource consumption ratio and the task processing difficulty.

7. The method according to claim 1, characterized in that After determining a target number of inference layers from a plurality of model layers in the inference model according to the importance parameter, and using the index of the inference layer as a preset index set corresponding to the preset task identifier, the method further includes: For each inference layer, insert the initial dimension reduction matrix and the initial dimension increase matrix into the initial parameter matrix corresponding to the inference layer; Fine-tune the inference model using a second task data set corresponding to a preset task identifier, and perform gradient update on the initial dimension reduction matrix and the initial dimension increase matrix to obtain a layer dimension reduction matrix and a layer dimension increase matrix corresponding to the inference layer; The layer index corresponding to the reasoning layer is used as a preset index, and the layer dimension reduction matrix and the layer dimension increase matrix are used as layer task matrices; The layer task matrices corresponding to all preset indexes are used as preset task matrices, and an association relationship between the preset task matrix and the preset task identifier is established.

8. A task-based model reasoning device, the device being used to implement the method according to any one of claims 1 to 7, characterized in that: The device comprises: A data acquisition module, used to acquire data to be processed and a task identifier corresponding to the data to be processed, wherein the data to be processed is one or more of text, image, and video; An index acquisition module, used to acquire an index set corresponding to the data to be processed by using a first association relationship between a preset task identifier and a preset index set and the task identifier, wherein the index set includes a plurality of preset indexes; A matrix acquisition module, used to acquire a task matrix corresponding to the data to be processed by using the task identifier, wherein the task matrix includes a layer task matrix associated with the preset index; An inference module, configured to transform the data to be processed using a layer task matrix corresponding to all inference layers in the inference model to obtain an inference result of the data to be processed, wherein the layer index of the inference layer in the inference model is the preset index; The first association relationship is obtained through the following steps: Obtain a first task data set corresponding to a preset task identifier; perform a first reasoning on the first task data set using all model layers in the reasoning model to obtain a baseline performance parameter of the reasoning model and a layer input vector of each model layer; for each model layer in the reasoning model, perform a second reasoning on the first task data set using other model layers except the model layer to obtain a layer performance parameter corresponding to the model layer; use the baseline performance parameter and the layer performance parameter as the specified performance parameters of the reasoning model; use the specified performance parameter and the layer input vector of each model layer to calculate the importance parameter of each model layer under the first task data set; calculate a target number based on the layer input vector of each model layer and the reasoning resource consumption; determine a target number of reasoning layers from multiple model layers in the reasoning model according to the importance parameter, and use the index of the reasoning layer as the preset index set corresponding to the preset task identifier.

Citation Information

Patent Citations

  • Processing method and device of reasoning service, computer equipment and storage medium

    CN115358401A

  • Speech recognition fine tuning task acceleration method based on low-rank matrix approximation

    CN117059103A

  • Method for acquiring models and computing device

    WO2024217135A1