Model processing method and device and storage medium
By training network parameters and architecture parameters on the submodules of the pre-trained model, and optimizing model output using candidate operations, the high cost problem of algorithm updates in artificial intelligence IoT devices is solved, and efficient model optimization and resource utilization are achieved.
Patent Information
- Application Number
- CN202410046359.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, in artificial intelligence IoT devices, there are large development and maintenance costs when algorithm updates, and it is impossible to effectively coordinate the collaboration and resource requirements between different modules, resulting in difficulty in performance optimization.
By training network parameters and architectural parameters of the pretrained model submodules, using candidate operations such as freezing, fine-tuning, and inserting adapters, optimize model output, simplify the training process and reduce manual intervention.
It realizes that without reducing model performance, the training process is simplified, the model adaptability and resource utilization efficiency are improved, and the development cost is reduced.
Smart Images

Figure CN120296407A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of algorithm optimization, model fine-tuning training, and artificial intelligence, and particularly relates to a model processing method, apparatus, and storage medium. Background Art
[0002] With the progress of technology and the development of the field of artificial intelligence, people's requirements for the functions of electronic devices used in daily life are constantly increasing. Therefore, relevant technical personnel need to continuously develop or update the corresponding functions of electronic devices to meet the gradually increasing usage requirements of users.
[0003] Currently, there are many functions implemented based on deep learning algorithms in artificial intelligence Internet of Things devices. For example, functions such as gesture recognition, human portrait tracking, and voice interaction. However, when these functions need to be upgraded, how to avoid algorithm updates and improve the reuse rate of the implemented algorithms is an urgent problem for relevant technical personnel to solve. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a model processing method, apparatus, and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a model processing method, including: in response to a second device receiving a training request sent by a first device, determining a pre-trained model, where the pre-trained model includes one or more sub-modules, and the training request is used to request the second device to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model, the network parameters are used to adjust the functions corresponding to one or more sub-modules in the pre-trained model, and the architecture parameters are used to adjust the connection paths between sub-modules in the pre-trained model; processing the pre-trained model in the following manner to obtain a trained model and sending it to the first device: based on the pre-trained model, training the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model, where the network parameters are parameters for performing candidate operations on sub-modules, and the architecture parameters are parameters for sub-modules to select candidate paths; based on the network parameters and / or the architecture parameters, determining the output of the pre-trained model.
[0006] In one implementation, the determining the output of the pre-trained model based on the network parameters and / or the architecture parameters includes: determining the network parameters corresponding to candidate operations of one or more sub-modules; and / or determining the architecture parameters of the paths between one or more sub-modules; performing a weighted sum on the network parameters and / or the architecture parameters to determine the output of the pre-trained model.
[0007] In one implementation, the candidate operations include at least one of the following: a freezing operation, a fine-tuning operation, and inserting an adapter; the weighted summation of the network parameters and / or the architecture parameters to determine the output of the pre-trained model includes: determining the sub-module network parameters as the network parameters corresponding to the fine-tuning operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; or determining the sub-module network parameters as the network parameters corresponding to the freezing operation, and determining the network parameters corresponding to the inserted adapter of the sub-module, as well as the architecture parameters corresponding to the path of the inserted adapter, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; or determining the sub-module network parameters as the network parameters corresponding to the freezing operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model; summing the outputs of the one or more sub-modules to determine the output of the pre-trained model.
[0008] In one implementation, training the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model includes: training the network parameters based on a training set; training the architecture parameters based on a validation set, where the training set and the validation set are composed of different training data.
[0009] In one implementation, the method further includes: attaching the determined output of the pre-trained model to a loss function.
[0010] In one implementation, the loss function includes: taking the number of parameters of the architecture parameters trained for the one or more sub-modules as weights, and performing a weighted summation of the architecture parameters trained for the one or more sub-modules.
[0011] In one implementation, the second device is a cloud, and the first device is a terminal.
[0012] According to a second aspect of the embodiments of the present disclosure, there is provided a model processing apparatus, including:
[0013] A search unit, configured to determine a pre-trained model in response to a training request sent by a first device received by a second device, where the pre-trained model includes one or more sub-modules, the training request is used to request the second device to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model, the network parameters are used to adjust the functions corresponding to one or more sub-modules in the pre-trained model, and the architecture parameters are used to adjust the connection paths between sub-modules in the pre-trained model;
[0014] A sending unit, configured to process the pre-trained model through a training unit and a determining unit in the following manner to obtain a trained model, and send it to the first device:
[0015] The training unit is configured to train network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model based on the pre-trained model, where the network parameters are parameters for performing candidate operations on the sub-module, and the architecture parameters are parameters for the sub-module to select a candidate path;
[0016] The determining unit is configured to determine the output of the pre-trained model based on the network parameters and / or the architecture parameters.
[0017] In one implementation, the determining unit determines the output of the pre-trained model based on the network parameters and / or the architecture parameters in the following manner: determining network parameters corresponding to one or more sub-module candidate operations; and / or determining architecture parameters of paths between one or more sub-modules; performing a weighted sum on the network parameters and / or the architecture parameters to determine the output of the pre-trained model.
[0018] In one implementation, the candidate operations include at least one of the following: a freezing operation, a fine-tuning operation, and inserting an adapter; the determining unit performs a weighted sum on the network parameters and / or the architecture parameters in the following manner to determine the output of the pre-trained model: determining that the sub-module network parameter is the network parameter corresponding to the fine-tuning operation, and determining the architecture parameter corresponding to the sub-module, and taking the value obtained by multiplying the network parameter and the weighted architecture parameter as the output of the sub-module; or determining that the sub-module network parameter is the network parameter corresponding to the freezing operation, and determining the network parameter corresponding to the adapter inserted into the sub-module, and the architecture parameter of the path where the adapter is inserted, and taking the value obtained by multiplying the network parameter and the weighted architecture parameter as the output of the sub-module; or determining that the sub-module network parameter is the network parameter corresponding to the freezing operation, and determining the architecture parameter corresponding to the sub-module, and taking the value obtained by multiplying the network parameter and the weighted architecture parameter as the output of the pre-trained model; summing the outputs of the one or more sub-modules to determine the output of the pre-trained model.
[0019] In one implementation, the training unit trains the network parameters and / or the architecture parameters of one or more sub-modules in the pre-trained model in the following manner: training the network parameters based on a training set; training the architecture parameters based on a validation set, where the training set and the validation set are composed of different training data.
[0020] In one implementation, the device further includes: a control unit, configured to attach the output to a loss function based on the determined output of the pre-trained model.
[0021] In one implementation, the loss function includes: performing a weighted sum of the architecture parameters trained for the one or more sub-modules based on the number of parameters of the architecture parameters trained for the one or more sub-modules.
[0022] In one implementation, the second device is a cloud and the first device is a terminal.
[0023] According to a third aspect of the embodiments of the present disclosure, there is provided a model processing device, including:
[0024] a processor;
[0025] a memory for storing instructions executable by the processor;
[0026] wherein the processor is configured to: execute the method described in the first aspect or any one of the implementations of the first aspect.
[0027] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions that, when executed by a processor of a terminal, enable the terminal to execute the model processing method described in the first aspect or any one of the implementations of the first aspect.
[0028] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The second device trains the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model, determines the output of the pre-trained model based on the network parameters and / or architecture parameters, and sends the trained model to the first device. By training the network parameters and / or architecture parameters of the pre-trained model, the training process is simplified and manual intervention is reduced, thereby effectively optimizing the pre-trained model without degrading its performance.
[0029] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0031] Figure 1 is a flowchart of a model processing method shown according to an exemplary embodiment.
[0032] Figure 2It is a flowchart of a model processing method shown according to an exemplary embodiment.
[0033] Figure 3 It is a schematic diagram of a cascaded model processing method shown according to an exemplary embodiment.
[0034] Figure 4 It is a flowchart of a model processing method shown according to an exemplary embodiment.
[0035] Figure 5 It is a flowchart of a model processing method shown according to an exemplary embodiment.
[0036] Figure 6 It is a flowchart of a model processing method shown according to an exemplary embodiment.
[0037] Figure 7 It is a schematic diagram of the interaction between the cloud and the terminal in a model processing method shown according to an exemplary embodiment.
[0038] Figure 8 It is a block diagram of a model processing device shown according to an exemplary embodiment.
[0039] Figure 9 It is a block diagram of a device for model processing shown according to an exemplary embodiment. Detailed implementation
[0040] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.
[0041] With the progress of technology and the development of the field of artificial intelligence, people's requirements for the functions of daily electronic devices are constantly increasing, and the functions implemented based on deep learning algorithms are also used more and more. For example, voice interaction, gesture recognition in artificial intelligence Internet of Things devices, and portrait tracking and person recognition in monitoring devices are all directly or indirectly implemented by using deep learning models. However, after the relevant functions are run, operations such as updating, optimizing, or function iteration of the algorithm or model are required. However, in the related technology, the update of the algorithm generally replaces the entire algorithm completely. Therefore, there are relatively large costs, including development, maintenance, testing, and the product life cycle. In order to improve the advantages of algorithm reuse and avoid relatively large development costs, the model fine-tuning technology in the related technology can be used for better optimization and update iteration of the algorithm.
[0042] In related technologies, model fine-tuning technology is an important research direction in the field of deep learning, and there have been many related research works. In terms of model fine-tuning, the main focus is on how to better utilize the knowledge of pre-trained models to improve the effect of fine-tuning. For example, some research works explore how to select better pre-trained models, how to adjust the hyperparameters of fine-tuning, how to design better fine-tuning strategies, and so on. In terms of multi-task learning, the main focus is on how to design better multi-task learning frameworks to improve the generalization ability and effect of the model. For example, some research works explore how to select better task combinations, how to design better shared layer structures, how to adjust the weights of different tasks, and so on. In terms of domain adaptation, the main focus is on how to better utilize the knowledge of pre-trained models to adapt to the data of new domains. For example, some research works explore how to select better pre-trained models, how to design better domain adaptation strategies, how to utilize unlabeled data, and so on. In terms of knowledge distillation, the main focus is on how to better utilize the knowledge of the teacher model to guide the training of the student model. For example, some research works explore how to select better teacher models, how to design better knowledge distillation strategies, how to utilize different types of knowledge, and so on. In terms of adversarial training, the main focus is on how to better utilize adversarial samples to improve the robustness and generalization ability of the model. For example, some research works explore how to design better adversarial sample generation algorithms, how to use adversarial samples for training, how to adjust the hyperparameters of adversarial training, and so on. In terms of self-supervised learning, the main focus is on how to better utilize unlabeled data to pre-train the model. For example, some research works explore how to design better self-supervised tasks, how to utilize different types of unlabeled data, how to adjust the hyperparameters of pre-training, and so on. In summary, the research status of model fine-tuning technology is very active, and researchers are constantly exploring new methods and technologies to improve the performance and generalization ability of the model. The above technologies can be used alone or in combination to improve the performance and generalization ability of the model.
[0043] In addition, for an algorithm system composed of multiple series / parallel single-task models, where the output of the previous model or the intermediate layer is passed as input to the latter, or the parallel results of several modules are passed to the subsequent module, the cascade structure enables each task to utilize the existing trained models and data resources, while enabling a good match between upstream and downstream tasks. However, transferring the cascade model to a specific domain is data-consuming and time-consuming. To better adapt the trained algorithm to the deployed environment, such as mobile devices, personal computers (PCs), other system-on-chips (SoCs), and to adapt to different application environments, fine-tuning the algorithm using actual data becomes an important task. However, the cost of fine-tuning the entire algorithm framework is too high. Usually, the performance adaptation function is achieved by fine-tuning some structural parameters of the model, or for multi-modal tasks, by fine-tuning some tasks or the relevant model structures of some tasks. To achieve cost-effective transmission, various adapters have been proposed in the related art, which contain a small number of model parameters. Although these methods focus on the compression and expansion of pre-trained models within different adapter modules, little research has been done on the adapter insertion strategy for different tasks.
[0044] However, in the related art, the parameters and memory efficiency of fine-tuning the complete cascade model are low, and applying only the adapter module to the cascade model cannot achieve performance comparable to that of fine-tuning. When facing a single task, in the related art, the cooperation between different tuning algorithms or modules is ignored. In the research of cascade models, the adaptation of the upstream pre-trained model often cooperates with the fine-tuning of the downstream task, making a trade-off between resource requirements and the final result. In some related algorithm tuning schemes, the context model cannot be coordinated and considered. For multi-task cascade training, only the coordination between multi-tasks is considered, but the effectiveness of each single task itself has not been explored through experiments, and the number of trainable parameters required for its own function and downstream tasks cannot be reduced. Moreover, for the related methods of adaptive adapters, the insertion position cannot achieve performance similar to that of full transmission. The adapters at higher layers have a greater impact than those at lower layers, and full fine-tuning is not as good as only fine-tuning the top layer.
[0045] Therefore, in the embodiments of the present disclosure, a technical solution is provided: The second device trains the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model, determines the output of the pre-trained model based on the network parameters and / or architecture parameters, and sends the trained model to the first device. By training the network parameters and / or architecture parameters of the pre-trained model, the training process is simplified and manual intervention is reduced, thereby effectively optimizing the pre-trained model without degrading its performance.
[0046] Figure 1It is a flowchart of a model processing method shown according to an exemplary embodiment. As Figure 1 shown, it includes the following steps.
[0047] In step S11, in response to the second device receiving a training request sent by the first device, a pre-trained model is determined.
[0048] Among them, the pre-trained model includes one or more sub-modules.
[0049] In the embodiments of the present disclosure, the second device can search for a pre-trained model based on a search space, for example, a big data space, so as to use the optimal architecture of the search space as the pre-trained model. Among them, the optimal architecture can be the model architecture with the most perfect functions, or it can also be the model architecture with the most stable functions.
[0050] In the embodiments of the present disclosure, the pre-trained model includes a single-task model or a cascade model. Among them, the single-task model includes one sub-module, and the cascade model includes multiple models, where each model contains multiple different sub-modules.
[0051] In the embodiments of the present disclosure, the training request can be a model training request corresponding to the function that needs to be updated after the first device determines the function that needs to be updated. Among them, the model corresponding to the function can be the model used to implement the corresponding function. And the training request can be used to specify the training of one or more sub-modules in the pre-trained model. Or the training request can specify the training of all sub-modules of the pre-trained model.
[0052] In the embodiments of the present disclosure, the pre-trained model can be located in the second device. The second device can determine the pre-trained model in the search space based on the training request sent by the first device. Among them, the pre-trained model can correspond to the system version and hardware structure of the first device. For example, for different electronic devices, the number of cameras can be different. Therefore, the models corresponding to the shooting function are also different. The cloud can determine, based on the update request, the pre-trained model that can implement the number of cameras of the electronic device and perform training. In the embodiments of the present disclosure, the second device can train any one of the network parameters or architecture parameters of the sub-module specified by the training request. Or the second device can simultaneously train the network parameters and architecture parameters of the sub-module specified by the training request. Among them, it can be understood that training the network parameters can optimize the function of the sub-module; training the architecture parameters can determine the connection path between sub-modules.
[0053] In the embodiments of the present disclosure, the pre-trained model can be a cascade architecture model. Or the pre-trained model can include a cascade architecture model. For example, the pre-trained model can be a task cascade architecture.
[0054] In step S12, the pre-trained model is processed in the following manner to obtain a trained model, which is then sent to the second device.
[0055] In step S13, based on the pre-trained model, the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model are trained.
[0056] Among them, the network parameters are the parameters for performing candidate operations on the sub-module, and the architecture parameters are the parameters for the sub-module to select candidate paths.
[0057] In the embodiments of the present disclosure, different pre-trained models may include sub-modules with different functions, and the overall connection frameworks of different pre-trained models are different, that is, the connection methods of the sub-modules in different pre-trained models may be different. Therefore, training the pre-trained model can be understood as adjusting the corresponding functions of the sub-modules of the model and determining the connection methods between the sub-modules. Among them, the connection method includes the connection relationship between the sub-modules. For example, the connection relationship between the sub-modules may be that sub-module A can be connected to sub-module B, or sub-module A can be connected to sub-module C. For different connection relationships, there may be different architecture parameters.
[0058] In step S14, based on the network parameters and / or architecture parameters, the output of the pre-trained model is determined.
[0059] In the embodiments of the present disclosure, the network space can be composed of the candidate operations performed by the sub-modules of the pre-trained model and the candidate paths of the sub-modules. Among them, the candidate operations performed by the sub-modules can be used as the network parameters of the pre-trained model, and the candidate paths selected by the sub-modules can be used as the architecture parameters of the pre-trained model.
[0060] In the embodiments of the present disclosure, the first device can determine the pre-trained model based on the update instruction of the entire terminal system. Or the first device can determine the pre-trained model based on the information feedback of the corresponding function of the model. For example, when it is detected that the gesture function fails multiple times, relevant technical personnel determine that the model corresponding to the gesture function needs to be pre-trained.
[0061] In the embodiments of the present disclosure, a search space can be constructed based on multiple models, including single-task models and / or cascade models, as well as candidate operations.
[0062] In the embodiments of the present disclosure, by training the network parameters and / or architecture parameters of one or more sub-modules in the determined pre-trained model, determining the output of the pre-trained model based on the network parameters and / or architecture parameters, and having the second device send the pre-trained model that has completed training to the first device, the training process is simplified and manual intervention is reduced, thereby effectively optimizing the pre-trained model without degrading its performance.
[0063] In the embodiments of the present disclosure, a single-task model may include one sub-module. A cascade model may include multiple models, where each model may contain one or more sub-modules.
[0064] In the embodiments of the present disclosure, each sub-module may contain network parameters and / or architecture parameters.
[0065] Figure 2 It is a flowchart of a model processing method shown according to an exemplary embodiment, as Figure 2 shown, and includes the following steps.
[0066] In step S21, determine the network parameters corresponding to one or more sub-module candidate operations; and / or determine the architecture parameters of the paths between one or more sub-modules.
[0067] In step S22, perform a weighted sum on the network parameters and / or architecture parameters to determine the output of the pre-trained model.
[0068] In the embodiments of the present disclosure, a sub-module may determine whether there are network parameters in subsequent sub-modules by automatically searching for a method to judge whether they are the network parameters of the current sub-module.
[0069] In the embodiments of the present disclosure, the architecture parameters may be set as path weights.
[0070] In the embodiments of the present disclosure, for a single-task model, there may be only one network parameter and / or one architecture parameter.
[0071] In the embodiments of the present disclosure, for different sub-modules corresponding to multiple models in a cascade model, each sub-module may have one network parameter and / or one architecture parameter. For a cascade model with multiple models, the output of the previous model may be used as the input of the next model.
[0072] In the embodiments of the present disclosure, Figure 3 It is a schematic diagram of a cascade model processing method shown according to an exemplary embodiment. A cascade model may include multiple tasks, that is, multiple models, where an adaptive fine-tuning module may be understood as a sub-module after corresponding pre-training of the sub-module. As Figure 3 shown, the cascade model includes Task 1, Task 2, Task 3 to Task M, where each task, that is, in the model, also corresponds to the same or different sub-modules. As Figure 3 shown, Task 1, and Model 1 contains two sub-modules, and the two sub-modules can be pre-trained to obtain pre-trained sub-module 1 and pre-trained sub-module 2 containing network parameters and / or architecture parameters. For the same reason, in Task 2, Task 3 to Task M, there is one or more pre-trained sub-modules containing corresponding network parameters and / or architecture parameters. And, as Figure 3As shown, the outputs of Task 1 and Module 1 can be used as the inputs of Task 2 and Module 2.
[0073] In the embodiments of the present disclosure, the network parameters and / or architecture parameters of each sub-module in one or more sub-modules can be weighted and summed to determine the output of the model.
[0074] In the embodiments of the present disclosure, by training the network parameters and / or architecture parameters of the pre-trained model, the training process is simplified and manual intervention is reduced, and thus the pre-trained model can be effectively optimized without degrading its performance.
[0075] In the disclosed embodiments, the functions of the sub-modules can be adjusted accordingly based on different candidate operations. Among them, the candidate operations include one of the following: freeze operation, fine-tuning operation, and inserting an adapter. Alternatively, the candidate operations can be any combination of the above operations.
[0076] Among them, the freeze operation can be not training the sub-module.
[0077] Among them, the fine-tuning operation can be re-training the module based on the existing module.
[0078] Among them, the sub-module can correspond to an adaptation search unit, and the adapter insertion operation can be performed after one or more sub-modules, and different adapters have different names.
[0079] In the embodiments of the present disclosure, the freeze operation can be preset by a person skilled in the art, or it can be determined whether to perform the freeze operation by judging whether the function corresponding to the sub-module is the same as before.
[0080] In the embodiments of the present disclosure, the candidate operations can be processed for different sub-modules, or can also be processed for the neural network.
[0081] Figure 4 is a flowchart of a model processing method shown according to an exemplary embodiment, as Figure 4 shown, including the following steps.
[0082] In step S31, determine the network parameters corresponding to one or more sub-module candidate operations; and / or determine the architecture parameters of the paths between one or more sub-modules.
[0083] In step S321, determine that the sub-module network parameters are the network parameters corresponding to the fine-tuning operation, and determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model.
[0084] In step S322, determine that the sub-module network parameters are the network parameters corresponding to the freezing operation, determine the network parameters corresponding to the sub-module insertion adapter, and the architecture parameters corresponding to the path of the insertion adapter, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model.
[0085] In step S323, determine that the sub-module network parameters are the network parameters corresponding to the freezing operation, determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model.
[0086] In step S33, sum the outputs of one or more sub-modules to determine the output of the pre-trained model.
[0087] In the embodiments of the present disclosure, the outputs of one or more sub-modules can be determined in different ways based on different candidate operations.
[0088] In the embodiments of the present disclosure, step S21 is the same as step S31, and will not be elaborated here.
[0089] In the embodiments of the present disclosure, when the candidate operation corresponding to the sub-module only includes the fine-tuning operation, determine that the sub-module network parameters are the network parameters corresponding to the fine-tuning operation, determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module.
[0090] In the embodiments of the present disclosure, when the candidate operation corresponding to the sub-module includes the freezing operation and it is determined that an adapter can be inserted, determine that the sub-module network parameters are the network parameters corresponding to the freezing operation, determine the network parameters corresponding to the sub-module insertion adapter, and the architecture parameters corresponding to the path of the insertion adapter, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module.
[0091] In the embodiments of the present disclosure, when the candidate operation corresponding to the sub-module includes the freezing operation and it is determined that no adapter is inserted, determine that the sub-module network parameters are the network parameters corresponding to the freezing operation, determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module.
[0092] In the embodiments of the present disclosure, since the pre-trained model may include one or more sub-models, the sum of the outputs of one or more sub-models can be used as the final output of the pre-trained model.
[0093] In the embodiments of the present disclosure, Figure 5 is a flowchart of a model processing method shown according to an exemplary embodiment. As Figure 5As shown, each path can represent a calculation process for determining the outputs of different sub - modules. Among them, in Path 1, the candidate operation of the sub - module is the fine - tuning operation. Determine the network parameters of the fine - tuning operation, and determine the path weight corresponding to the sub - module as the architecture parameter. Multiply the network parameters by the architecture parameter as the output of the sub - module. In Path 2, the candidate operation of the sub - module is the freezing operation. Determine the network parameters of the freezing operation, and determine the path weight corresponding to the sub - module as the architecture parameter. Multiply the network parameters by the architecture parameter as the output of the sub - module. In Path 2, the candidate operation of the sub - module is the freezing operation. Determine the network parameters of the freezing operation, and determine that the sub - module performs the operation of inserting an adapter after the freezing operation. Determine the network parameters corresponding to the adapter, and determine the path weight corresponding to the sub - module as the architecture parameter. Multiply the network parameters by the architecture parameter as the output of the sub - module. And finally, the sum of the outputs of multiple sub - modules is used as the final output.
[0094] In the embodiments of the present disclosure, by setting candidate operations, including fine - tuning operations, freezing operations, and inserting adapters, the extended fine - tuning operation enables the system to directly determine the final tuning strategy, replacing manual setting and cumbersome experiments. Moreover, the cooperation between the module and the adapter adjustment can achieve the best performance of the module. In addition, integrating the fine - tuning operation with the insertion of the adapter and taking the freezing operation as the candidate operation of each adapted sub - module simplifies the training process and reduces manual intervention.
[0095] In the embodiments of the present disclosure, the network parameters and architecture parameters can be trained based on training data.
[0096] Figure 6 is a flowchart of a model processing method shown according to an exemplary embodiment. As Figure 6 shown, it includes the following steps.
[0097] In step S41, train the network parameters and / or architecture parameters of one or more sub - modules in the pre - trained model.
[0098] In step S421, based on the training set, train the network parameters.
[0099] In step S422, based on the validation set, train the architecture parameters.
[0100] Among them, the training set and the validation set are composed of different training data.
[0101] In the embodiments of the present disclosure, determine the architecture parameters through the validation set while freezing the network parameters.
[0102] In the embodiments of the present disclosure, determine the adaptation or fine - tuning parameters through the training set, and the architecture parameters remain unchanged.
[0103] In the embodiments of the present disclosure, the hard architecture weights of each sub-module can be calculated. Among them, the hard architecture weights can be the architecture parameters of the sub-module based on divisors. For example, the architecture weights are restricted between 0 and 1, so as to avoid inconsistencies between the evaluation and training processes. Among them, during the training process, the Softmax function can be used to calculate the architecture parameters of each path, so as to achieve path sampling without truncating the gradient.
[0104] In the embodiments of the present disclosure, in order to control the size of the learning structure, the architecture learning can be guided by adding a penalty to the loss.
[0105] In the embodiments of the present disclosure, based on the output of the determined pre-trained model, the output is appended to the loss function.
[0106] Among them, the loss function includes: taking the number of parameters of the architecture parameters trained by one or more sub-modules as weights, and performing a weighted sum of the architecture parameters trained by one or more sub-modules.
[0107] In the embodiments of the present disclosure, the loss function can be determined by the following formula:
[0108]
[0109] Loss is a proposed loss penalty term related to the number of trainable parameters and architecture parameters, where i is the module index, N is the total number of modules in the cascaded model, x i,ft , x i,ad and x i,fr are the architecture parameters of the fine-tuning operation, freezing operation, and inserting adapter respectively. P i,ft , P i,ad and P i,fr are the number of trainable parameters for each operation, where P i,ft is equal to the size of the module, P i,ad is the number of parameters of the adapter, and P i,fr can be set to 0 or a larger number to reduce its order.
[0110] Among them, by setting P i,fr to half of P i,fr , that is, the middle of P i,ad and P i,fr , and dynamically changing with different modules. Thus, it is avoided that P i,fr is set to 0 or a decimal less than P i,ad , resulting in the final search result converging to an architecture with a large frozen part, leading to a non-optimal state of the model performance. In addition, the weighted sum term is normalized by the sum of the number of trainable parameters of the three operations to ensure that the parameter penalty term and the original loss are in a similar range.
[0111] In the embodiments of the present disclosure, the number of parameters that need to be adjusted or fine-tuned on each path is measured, multiplied by the architecture parameters on that path, and finally added and directly appended to the original loss function, which can promote more module selection to insert adapters or freeze operations. It can be understood that in the embodiments of the present disclosure, the architecture of the search is compressed by parameter penalty of the loss, so as to achieve optimization in an adaptive manner and reduce the influence of human factors.
[0112] In the embodiments of the present disclosure, to avoid the normal use of a terminal device being affected by the model processing process, the architecture of the mobile end and the cloud data can be synchronized by establishing cloud optimization, so as to realize the update of the terminal device model.
[0113] In the embodiments of the present disclosure, in response to the cloud receiving a training request sent by a first device, a pre-trained model is determined. The pre-trained model includes one or more sub-modules. The pre-trained model is processed to obtain a trained model and sent to the terminal.
[0114] Figure 7 is a schematic diagram of the interaction between the cloud and the terminal in a model processing method shown according to an exemplary embodiment, as Figure 7 shown. The terminal determines different models corresponding to the task and sends a training request for the task to the cloud. The cloud trains the model corresponding to the task based on the training request for the task sent by the terminal and feeds back the trained model to the terminal.
[0115] In the implementation of the present disclosure, after the optimization process of the model is completed in the cloud, some modules or certain functional layers or a certain task are sent down to update the mobile algorithm system. After the adaptive training of the model is completed, first perform cloud adaptation, and automatically send the completed adapted part to the end side for update, and the actual evaluation results can be monitored in real time to facilitate subsequent adjustment work.
[0116] In the embodiments of the present disclosure, a cascaded model combining the three tasks of speech enhancement (SE), automatic speech recognition (ASR), and natural language processing (NLP) is used to illustrate the model processing method.
[0117] In the embodiments of the present disclosure, experiments are conducted based on the open-source corpus (Simple Language Understanding and Reasoning Platform, SLURP), which includes speech data, ASR tags, and semantic tags. It contains 85 hours of training set, 7 hours of validation set, and 10 hours of test set. The SE and ASR models are pre-trained based on different databases, including Voicebank and Librispeech, and the NLP model is randomly initialized.
[0118] In the embodiments of the present disclosure, the SE model consists of 6 Transformer blocks as encoders, each block having 4 heads, 256 channels, and 128 hidden nodes. The candidate adapter of the SE model is inserted at the end of each Transformer block.
[0119] The ASR model uses CRDNN (a CNN, a RNN, and a MLP) as the encoder, including 3 convolutional neural network CNN blocks and 3 deep neural network DNN blocks. Each CNN block contains two layers of CNN. Possible adapters are added after each CNN layer, and the adapters in the DNN block are inserted after the linear projection layer. The decoders of ASR and NLP are both Attentional-RNN, which contains 1 RNN block and 1 Linear layer. The candidate adapters in the ASR and NLP decoders are added after the RNN block and the Linear layer. The output Embedding after Softmax in the ASR decoder is directly sent to the NLP decoder as the input for gradient backpropagation.
[0120] In the embodiments of the present disclosure, the loss only considers the NLP loss (Softmax + Cross-Entropy) and the parameter penalty term Loss. The adaptive adapter is selected as the candidate operation, which consists of an encoder, a non-linear layer, and a decoder, and has skip connections. The size of its hidden layer is set to one-fourth of the input size. It is optimized by performing a first-order approximation of the differentiable architecture search (DARTS) algorithm on the adaptive adapter. The training data is divided into two equal parts, one part for architecture parameter learning and the other part for network parameter training. The learning rate for fine-tuning and network parameter learning can be set to 0.0001, while the learning rate for architecture parameter learning can be set to 0.001.
[0121] In the embodiments of the present disclosure, full finetuning, which fine-tunes all parameters in the cascaded model, is a strong baseline for adaptation. However, to achieve the performance of full finetuning with fewer parameters, a commonly used parameter-efficient alternative strategy is to insert adapters in each module and fine-tune the downstream tasks. Due to the fine-tuning operation, the proposed adaptive fine-tuning module expands more trainable parameters. When the architecture is trained for several epochs, the finally selected parameters will drop sharply, especially after adding parameter control to the loss.
[0122] Furthermore, adopting a two-stage training method, where the architecture and network parameters are alternately updated in the first stage and only the network parameters are updated in the second stage, the performance is further enhanced and even exceeds full finetuning. The search architecture using less training data is the same as that of the full training set search, which can improve the training speed, obtain the best architecture with less data, and use the full data to update the network parameters.
[0123] Based on the same concept, the embodiments of the present disclosure also provide a model processing device.
[0124] It can be understood that, in order to implement the above functions, the model processing device provided by the embodiments of the present disclosure includes the corresponding hardware structure and / or software module for executing each function. Combining the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0125] Figure 8 It is a block diagram of a model processing device shown according to an exemplary embodiment. Referring to Figure 8 FIG., the device 100 includes a search unit 101, a sending unit 102, a training unit 103, a determination unit 104, and a control unit 105.
[0126] The search unit 101 is configured to determine a pre-trained model in response to the second device receiving a training request sent by the first device. The pre-trained model includes one or more sub-modules. The training request is used to request the second device to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model. The network parameters are used to adjust the functions corresponding to one or more sub-modules in the pre-trained model, and the architecture parameters are used to adjust the connection paths between the sub-modules in the pre-trained model.
[0127] The sending unit 102 is configured to process the pre-trained model through the training unit and the determination unit in the following manner to obtain a trained model and send it to the first device.
[0128] The training unit 103 is configured to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model based on the pre-trained model, where the network parameters are the parameters for performing candidate operations on the sub-module, and the architecture parameters are the parameters for the sub-module to select candidate paths.
[0129] The determination unit 104 is configured to determine the output of the pre-trained model based on the network parameters and / or architecture parameters.
[0130] The control unit 105 is configured to attach the output to the loss function based on the determined output of the pre-trained model.
[0131] In one implementation, the determination unit 104 determines the output of the pre-trained model based on the network parameters and / or architecture parameters in the following manner: determining the network parameters corresponding to one or more candidate operations of the sub-module, and / or determining the architecture parameters of the paths between one or more sub-modules, and performing a weighted sum on the network parameters and / or architecture parameters to determine the output of the pre-trained model.
[0132] In one implementation, the candidate operations include at least one of the following: freezing operation, fine-tuning operation, and inserting an adapter. The determination unit 104 performs a weighted sum on the network parameters and / or architecture parameters in the following manner to determine the output of the pre-trained model: determining that the network parameters of the sub-module are the network parameters corresponding to the fine-tuning operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module. Or determining that the network parameters of the sub-module are the network parameters corresponding to the freezing operation, and determining the network parameters corresponding to the inserted adapter of the sub-module, and the architecture parameters of the path where the adapter is inserted, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module. Or determining that the network parameters of the sub-module are the network parameters corresponding to the freezing operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model, and summing the outputs of one or more sub-modules to determine the output of the pre-trained model.
[0133] In one implementation, the training unit 103 trains the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model in the following manner: training the network parameters based on the training set, and training the architecture parameters based on the validation set, where the training set and the validation set are composed of different training data.
[0134] In one implementation, the loss function includes: performing a weighted sum of the architecture parameters trained for one or more sub-modules based on the number of parameters of the architecture parameters trained for one or more sub-modules.
[0135] In one implementation, the second device is the cloud and the first device is the terminal.
[0136] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0137] Figure 9 FIG. 200 is a block diagram of a device 200 for model processing according to an exemplary embodiment. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0138] Referring to Figure 9 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0139] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0140] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0141] The power component 206 provides power for various components of the device 200. The power component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.
[0142] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0143] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0144] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0145] The sensor assembly 214 includes one or more sensors for providing a status assessment of various aspects of the device 200. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0146] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0147] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0148] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the above instructions can be executed by a processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0149] It can be understood that the term "plural" in the present disclosure means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms of "a", "the", and "said" are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0150] It can be further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other, and do not represent a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0151] It can be further understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0152] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0153] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be advantageous.
[0154] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0155] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A model processing method, characterized in that, Including: In response to the second device receiving a training request sent by the first device, determining a pre-trained model, where the pre-trained model includes one or more sub-modules, and the training request is used to request the second device to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model. The network parameters are used to adjust the functions corresponding to one or more sub-modules in the pre-trained model, and the architecture parameters are used to adjust the connection paths between the sub-modules in the pre-trained model; Processing the pre-trained model in the following manner to obtain a trained model and sending it to the first device: Based on the pre-trained model, training the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model. Among them, the network parameters are the parameters for candidate operations on the sub-module, and the architecture parameters are the parameters for the sub-module to select candidate paths; Based on the network parameters and / or the architecture parameters, determining the output of the pre-trained model.
2. The method according to claim 1, characterized in that The determining the output of the pre-trained model based on the network parameters and / or the architecture parameters includes: Determining the network parameters corresponding to candidate operations of one or more sub-modules; and / or Determining the architecture parameters of the paths between one or more sub-modules; Performing a weighted sum on the network parameters and / or the architecture parameters to determine the output of the pre-trained model.
3. The method according to claim 2, wherein The candidate operations include at least one of the following: freeze operation, fine-tuning operation, and inserting an adapter; The performing a weighted sum on the network parameters and / or the architecture parameters to determine the output of the pre-trained model includes: Determining that the sub-module network parameters are the network parameters corresponding to the fine-tuning operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; or Determining that the sub-module network parameters are the network parameters corresponding to the freeze operation, and determining the network parameters corresponding to the inserted adapter of the sub-module, and the architecture parameters of the path corresponding to the inserted adapter, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; or Determining that the sub-module network parameters are the network parameters corresponding to the freeze operation, and determining the architecture parameters corresponding to the sub-module, and taking the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; Performing a sum on the outputs of the one or more sub-modules to determine the output of the pre-trained model.
4. The method according to claim 1, wherein The training the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model includes: Based on the training set, training the network parameters; Based on the validation set, training the architecture parameters, where the training set and the validation set are composed of different training data.
5. The method according to claim 1, characterized in that, The method further includes: Based on the determined output of the pre-trained model, attaching the output to a loss function.
6. The method according to claim 5, wherein The loss function includes: taking the number of parameters of the architecture parameters trained for the one or more sub-modules as weights, and performing a weighted sum on the architecture parameters trained for the one or more sub-modules.
7. The method according to claim 1, characterized in that The first device is a terminal, and the second device is a cloud.
8. A model processing device, characterized in that, Including: A search unit, configured to determine a pre-trained model in response to a second device receiving a training request sent by a first device, where the pre-trained model includes one or more sub-modules, and the training request is used to request the second device to train network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model. The network parameters are used to adjust the functions corresponding to one or more sub-modules in the pre-trained model, and the architecture parameters are used to adjust the connection paths between sub-modules in the pre-trained model; A sending unit, configured to process the pre-trained model in the following manner through a training unit and a determining unit to obtain a trained model, and send it to the first device: The training unit is configured to train the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model based on the pre-trained model, where the network parameters are parameters for performing candidate operations on the sub-modules, and the architecture parameters are parameters for the sub-modules to select candidate paths; The determining unit is configured to determine the output of the pre-trained model based on the network parameters and / or the architecture parameters.
9. The device according to claim 8, characterized in that, The determining unit determines the output of the pre-trained model based on the network parameters and / or the architecture parameters in the following manner: Determine the network parameters corresponding to candidate operations of one or more sub-modules; and / or Determine the architecture parameters of the paths between one or more sub-modules; Perform a weighted sum on the network parameters and / or the architecture parameters to determine the output of the pre-trained model.
10. The device according to claim 9, characterized in that, The candidate operations include at least one of the following: a freezing operation, a fine-tuning operation, and inserting an adapter; The determining unit performs a weighted sum on the network parameters and / or the architecture parameters in the following manner to determine the output of the pre-trained model: Determine that the network parameters of the sub-module are the network parameters corresponding to the fine-tuning operation, and determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; Or Determine that the network parameters of the sub-module are the network parameters corresponding to the freezing operation, and determine the network parameters corresponding to the adapter inserted into the sub-module, and the architecture parameters of the path where the adapter is inserted, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the sub-module; Or Determine that the network parameters of the sub-module are the network parameters corresponding to the freezing operation, and determine the architecture parameters corresponding to the sub-module, and use the value obtained by multiplying the network parameters and the weighted architecture parameters as the output of the pre-trained model; Sum the outputs of the one or more sub-modules to determine the output of the pre-trained model.
11. The device according to claim 8, characterized in that, The training unit trains the network parameters and / or architecture parameters of one or more sub-modules in the pre-trained model in the following manner: Train the network parameters based on a training set; Train the architecture parameters based on a validation set, where the training set and the validation set are composed of different training data.
12. The device according to claim 8, wherein, The apparatus further includes: A control unit, configured to attach the determined output of the pre-trained model to a loss function.
13. The device according to claim 12, wherein The loss function includes: taking the number of parameters of the architecture parameters trained by the one or more sub-modules as weights, and performing a weighted sum of the architecture parameters trained by the one or more sub-modules.
14. The device according to claim 8, characterized in that, The second device is a cloud, and the first device is a terminal.
15. A model processing device, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the model processing method according to any one of claims 1-7.
16. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to execute the method according to any one of claims 1-7.