Model processing method and device, model training method and device, electronic equipment, readable storage medium and program product
By loading shared branches and independent branches of the model to be run in the electronic device memory, the inefficiency problem in traditional model processing methods is solved, and faster model operation and task completion are achieved.
Patent Information
- Application Number
- CN202510330934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
In traditional model processing methods, there is a problem of low operational efficiency when multiple artificial intelligence models are processed in concert.
Load the shared target shared branches into electronic device memory, and determine the target model to be run from multiple target models, load its unique target independent branches for parallel or serial operation until all models are shared and independent branches are completed to achieve the target task.
It improves the running efficiency of the model, reduces memory usage, avoids device lag, improves device fluency, and maintains the accuracy of the model.
Smart Images

Figure CN120256053A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of imaging (Camera Technology), and particularly to a model processing method and apparatus, a model training method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the development of artificial intelligence technology, more and more electronic devices can deploy artificial intelligence models, such as AI (Artificial Intelligence) mobile phones, AI watches, AI glasses, AI bracelets, etc. Through the electronic devices deployed with artificial intelligence models, users' tasks (functions) can be better achieved. As tasks or functions become more and more complex, a task or function often requires at least two artificial intelligence models to cooperate in processing.
[0003] In traditional model processing methods, each model is usually deployed and run sequentially, resulting in low running efficiency. Summary of the Invention
[0004] Embodiments of this application provide a model processing method and apparatus, a model training method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the running efficiency of the model.
[0005] In a first aspect, this application provides a model processing method, which is applied to an electronic device. The electronic device includes at least two target models. The at least two target models include a common target shared branch and a target independent branch unique to each model. The method includes:
[0006] Loading the target shared branch into the memory of the electronic device;
[0007] Determining a target model to be run from the at least two target models, and loading the target independent branch of the target model to be run into the memory;
[0008] Running the target shared branch and the target independent branch of the target model to be run in the memory, and returning to execute the step of determining the target model to be run from the at least two target models until the target shared branch and the target independent branches of the at least two target models are all run to achieve a target task.
[0009] In one embodiment, after running the target shared branch and the target independent branch of the target model to be run in the memory, it further includes:
[0010] Deleting the target independent branch in the memory.
[0011] In one embodiment, the number of parameters of the target shared branch is less than that of any one of the at least two target models.
[0012] In one embodiment, before loading the target shared branch into the memory of the electronic device, the following steps are further included:
[0013] Training the to-be-trained shared branch shared by at least two to-be-trained models to obtain the trained target shared branch;
[0014] Fixing the parameters of the target shared branch, and training the to-be-trained independent branch unique to each to-be-trained model to obtain the target independent branch;
[0015] Based on the target shared branch and the target independent branch, obtaining at least two trained target models.
[0016] In one embodiment, the training accuracy of the to-be-trained independent branch is higher than that of the to-be-trained shared branch.
[0017] In a second aspect, the present application further provides a model processing device, which is applied to an electronic device. The electronic device includes at least two target models, and the at least two target models include a common target shared branch and a target independent branch unique to each model. The device includes:
[0018] A target shared branch loading module, configured to load the target shared branch into the memory of the electronic device;
[0019] An independent shared branch loading module, configured to determine a target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory;
[0020] A running module, configured to run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from the at least two target models until the target shared branch and the target independent branch of the at least two target models are all run to implement a target task.
[0021] In a third aspect, the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0022] Loading the target shared branch into the memory of the electronic device;
[0023] Determine the target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory;
[0024] Run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from the at least two target models until the target shared branches and target independent branches of the at least two target models are all run to achieve the target task.
[0025] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0026] Load the target shared branch into the memory of the electronic device;
[0027] Determine the target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory;
[0028] Run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from the at least two target models until the target shared branches and target independent branches of the at least two target models are all run to achieve the target task.
[0029] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0030] Load the target shared branch into the memory of the electronic device;
[0031] Determine the target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory;
[0032] Run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from the at least two target models until the target shared branches and target independent branches of the at least two target models are all run to achieve the target task.
[0033] The above model processing method, device, electronic device, computer-readable storage medium, and computer program product. The electronic device includes at least two target models. The at least two target models include a common target shared branch and a target independent branch unique to each model. The electronic device loads the target shared branch into the memory of the electronic device, determines the target model to be run from the at least two target models, and loads the target independent branch of the target model to be run into the memory. Then, the target shared branch and the target independent branch of the target model to be run can be run in the memory to implement the task corresponding to one of the target models, and the step of determining the target model to be run from the at least two target models is returned until the target shared branch and the target independent branches of the at least two target models are all run. That is to say, during the running of the at least two target models, the target shared branch is always loaded in the memory, and only the target independent branch of the target model to be run needs to be loaded into the memory to run the target model more quickly, thereby improving the running efficiency of the at least two target models and achieving the target task more quickly.
[0034] In a sixth aspect, the present application provides a model training method applied to an electronic device. The electronic device includes at least two models to be trained. The at least two models to be trained include a common shared branch to be trained and an independent branch to be trained unique to each model to be trained. The method includes:
[0035] Training the shared branch to be trained to obtain a trained target shared branch;
[0036] Fixing the parameters of the target shared branch, and for each independent branch to be trained, training the independent branch to be trained to obtain a target independent branch;
[0037] Based on the target shared branch and the target independent branch, obtaining at least two trained target models.
[0038] In a seventh aspect, the present application further provides a model training device applied to an electronic device. The electronic device includes at least two target models. The at least two target models include a common target shared branch and a target independent branch unique to each model. The device includes:
[0039] A shared branch training module for training the shared branch to be trained to obtain a trained target shared branch;
[0040] An independent branch training module for fixing the parameters of the target shared branch and, for each independent branch to be trained, training the independent branch to be trained to obtain a target independent branch;
[0041] A target model obtaining module, configured to obtain at least two trained target models based on the target shared branch and the target independent branch.
[0042] In an eighth aspect, the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0043] Train the to-be-trained shared branch to obtain a trained target shared branch;
[0044] Fix the parameters of the target shared branch, and for each to-be-trained independent branch, train the to-be-trained independent branch to obtain a target independent branch;
[0045] Based on the target shared branch and the target independent branch, obtain at least two trained target models.
[0046] In a ninth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0047] Train the to-be-trained shared branch to obtain a trained target shared branch;
[0048] Fix the parameters of the target shared branch, and for each to-be-trained independent branch, train the to-be-trained independent branch to obtain a target independent branch;
[0049] Based on the target shared branch and the target independent branch, obtain at least two trained target models.
[0050] In a tenth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0051] Train the to-be-trained shared branch to obtain a trained target shared branch;
[0052] Fix the parameters of the target shared branch, and for each to-be-trained independent branch, train the to-be-trained independent branch to obtain a target independent branch;
[0053] Based on the target shared branch and the target independent branch, obtain at least two trained target models.
[0054] The above model training method, device, electronic device, computer-readable storage medium and computer program product, the electronic device includes at least two target models, the at least two target models include a common target shared branch and a target independent branch unique to each model. Training the shared branch to be trained that is common to at least two models to be trained to obtain the trained target shared branch, and then fixing the parameters of the target shared branch, and training the independent branch to be trained unique to each model to be trained can effectively utilize the parameters of the target shared branch and effectively improve the model accuracy through the independent branches. Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a schematic flowchart of the model processing method in an embodiment;
[0057] Figure 2 It is a schematic diagram of the network structure of at least two target models in an embodiment;
[0058] Figure 3 It is a schematic diagram of the network structure of at least two target models in another embodiment;
[0059] Figure 4 It is a schematic diagram of the operation of at least two target models in an embodiment;
[0060] Figure 5 It is a schematic diagram of training the shared branch to be trained in an embodiment;
[0061] Figure 6 It is a schematic diagram of training the independent branch to be trained in an embodiment;
[0062] Figure 7 It is a schematic flowchart of the model training method in an embodiment;
[0063] Figure 8 It is a structural block diagram of the model processing device in an embodiment;
[0064] Figure 9 It is a structural block diagram of the model training device in an embodiment;
[0065] Figure 10 It is an internal structure diagram of the electronic device in an embodiment. Detailed Embodiments
[0066] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0067] In one embodiment, as Figure 1 shown, a model processing method is provided. In this embodiment, an example is given where this method is applied to an electronic device. The electronic device can be a terminal or a server. It can be understood that this method can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the electronic device includes at least two target models. The at least two target models include a common target sharing branch and a target independent branch unique to each model. The model processing method includes the following steps:
[0068] Step S102, load the target sharing branch into the memory of the electronic device.
[0069] Among them, the target model is an algorithm structure based on mathematical and computer science principles. By learning and analyzing a large amount of data, it can realize intelligent tasks such as predicting, classifying, and making decisions on unknown data. It is the core component of an artificial intelligence system, aiming to simulate human intelligent behavior and thinking patterns, so that the computer can automatically process and understand various complex information. The target model can be a neural network model. The neural network model is a mathematical model that imitates the structure and function of a biological neural network. It is composed of a large number of neurons connected to each other and can realize various artificial intelligence tasks, such as classification, regression, image recognition, speech recognition, etc., through learning and processing data.
[0070] The target shared branch is a network branch structure shared by at least two target models, that is, a network branch structure with the same parameters in at least two target models. The parameters of at least two target models may include weight parameters, and may also include bias parameters, activation function parameters, normalization layer parameters, or attention parameters, etc. The target independent branch is a network branch structure unique to each target model. Among at least two target models, the parameters of the target shared branch are the same, while the parameters of the target independent branch are different.
[0071] The target shared branch and the independent shared branch can run serially or in parallel. As Figure 2 shown, the electronic device includes two target models, and the two target models include target shared branch 1, target shared branch 2, and target shared branch 3. One of the target models includes a unique target independent branch 1, and the other target model includes a unique target independent branch 2, and target shared branch 2, target independent branch 1, and target independent branch 2 run in parallel. Due to the parallel design of the target independent branch and the target shared branch, the accuracy of the target model is effectively maintained, and the reduction in accuracy caused by model structure modification is avoided.
[0072] As Figure 3 shown, the electronic device includes two target models including target shared branch 1. One of the target models includes a unique target independent branch 1, and the other target model includes a unique target independent branch 2, and target shared branch 1 and the target independent branches (target independent branch 1 and target independent branch 2) run serially. The electronic device inputs input information into at least two target models, and processes the input information through at least two target models, enabling multiple target models to share model (weight) parameters, and at the same time, the accuracy of different subtasks can be guaranteed. In actual application processes, there can be more models for sharing model structures and weights.
[0073] Optionally, the electronic device trains at least two to-be-trained models to obtain at least two trained target models.
[0074] Optionally, the electronic device trains the to-be-trained shared branches in at least two to-be-trained models to obtain trained target shared branches; trains the to-be-trained independent branches in at least two to-be-trained models to obtain trained target independent branches.
[0075] Optionally, the number of parameters of the target shared branch is less than the number of parameters of any one of the at least two target models.
[0076] It can be understood that the number of parameters of the target shared branch is less than that of any one of at least two target models, that is, the memory occupancy of the target shared branch is less than that of any one of at least two target models. Therefore, during the operation of at least two target models, the target shared branch is continuously loaded into the memory with a smaller occupancy. Each target model to be run only needs to load the target independent branch into the memory to run the target model, which can improve the model operation efficiency with a smaller memory occupancy.
[0077] Step S104: Determine the target model to be run from at least two target models, and load the target independent branch of the target model to be run into the memory.
[0078] Optionally, the electronic device determines the target model to be run corresponding to the subtask to be implemented in the target task from at least two target models, and loads the target independent branch of the target model to be run into the memory.
[0079] For example, if the target task is to generate a video from the photos taken today, then first, the subtask to be implemented in this target task is to obtain the photos taken today. Then, determine the target model to be run corresponding to this subtask to be implemented from at least two target models. The target model to be run is the photo acquisition model. Next, the subtask to be implemented in this target task is to generate a video from the obtained photos. Then, determine the target model to be run corresponding to this subtask to be implemented from at least two target models. The target model to be run is the video generation model. Then, by running the photo acquisition model and the video generation model, it is possible to generate a video from the photos taken by the user today.
[0080] Step S106: Run the target shared branch and the target independent branch of the target model to be run in the memory, and return to the step of determining the target model to be run from at least two target models until the target shared branches and target independent branches of at least two target models are all run to implement the target task.
[0081] It can be understood that the target shared branch and the target independent branch of the target model to be run are loaded in the memory of the electronic device. Then, the target shared branch and the target independent branch constitute the target model to be run. The electronic device runs the target shared branch and the target independent branch of the target model to be run in the memory, that is, runs the target model in the memory to implement the subtask corresponding to the target model.
[0082] Exemplarily, the target task (target function) of the electronic device includes 4 target models. The sequence of operation of the target function is that after target model 1 finishes running, target model 2 runs, then target model 3 runs, and finally target model 4 runs to achieve the target task (target function).
[0083] Optionally, after running the target shared branch and the target independent branch of the target model to be run in the memory, it further includes: deleting the target independent branch in the memory.
[0084] After the electronic device runs the target shared branch and the target independent branch of the target model to be run in the memory, if the target model to be run has completed the corresponding subtask, the target independent branch in the memory is deleted. Then the target independent branch only occupies memory during operation, which can reduce the occupied space of the memory, and continue to run the target model to be run. Thus, during the operation of at least two target models, the occupied space of the required memory is reduced, computer resources are saved, and at the same time, the time required to load and delete the target independent branch is less, improving the efficiency of model operation.
[0085] As Figure 4 shown, the electronic device loads target shared branch 1, target shared branch 2, and target shared branch 3 into the memory, determines target model 1 to be run, and loads target independent branch 1 of target model 1 to be run into the memory, and runs target shared branch 1, target shared branch 2, target shared branch 3, and target independent branch 1 in the memory; after target shared branch 1, target shared branch 2, target shared branch 3, and target independent branch 1 run, delete target independent branch 1 in the memory; determine target model 2 to be run, and load target independent branch 2 of target model 2 to be run into the memory, and run target shared branch 1, target shared branch 2, target shared branch 3, and target independent branch 2 in the memory to achieve the target task.
[0086] In the above model processing method, the electronic device includes at least two target models. The at least two target models include a common target sharing branch and a target independent branch unique to each model. The electronic device loads the target sharing branch into the memory of the electronic device, determines the target model to be run from the at least two target models, and loads the target independent branch of the target model to be run into the memory. Then, the target sharing branch and the target independent branch of the target model to be run can be run in the memory to implement the task corresponding to one of the target models, and the step of determining the target model to be run from the at least two target models is returned until the target sharing branch and the target independent branches of the at least two target models are all run. That is to say, during the running of the at least two target models, the target sharing branch is always loaded in the memory, and only the target independent branch of the target model to be run needs to be loaded into the memory to run the target model more quickly, thereby improving the running efficiency of the at least two target models and achieving the target task more quickly.
[0087] Moreover, when the at least two target models are running, it can effectively reduce the memory occupancy of the electronic device, improve the fluency of the electronic device and avoid unfriendly situations such as the electronic device getting hot, effectively avoiding the lag and waiting time during the use of the mobile phone. The above model processing method reduces the problem of excessive memory when the target model is deployed, resulting in the failure of the previous item, improves the ability of the target model to be deployed, and at the same time, the above model processing method is applicable to various neural network models and has great generality.
[0088] In one embodiment, before loading the target sharing branch into the memory of the electronic device, it further includes: training the common to-be-trained sharing branch of the at least two to-be-trained models to obtain the trained target sharing branch; fixing the parameters of the target sharing branch, and training the to-be-trained independent branch for each to-be-trained model unique to obtain the target independent branch; obtaining at least two trained target models based on the target sharing branch and the target independent branch.
[0089] Optionally, the electronic device extracts the network branch structure with the same parameters from the at least two to-be-trained models as the to-be-trained sharing branch, and configures different to-be-trained independent branches for each to-be-trained model according to different subtasks. The independent branches of different tasks can have different structures, so while reducing the memory, the different independent branches can ensure the model accuracy corresponding to different tasks.
[0090] Optionally, the electronic device deletes the to-be-trained independent branches in the at least two to-be-trained models to obtain the common to-be-trained sharing branch of the at least two to-be-trained models; and trains the common to-be-trained sharing branch of the at least two to-be-trained models by using the loss function and labels of multiple tasks to obtain the trained target sharing branch.
[0091] The electronic device fixes the parameters of the target sharing branch, and trains the corresponding independent branches to be trained for different tasks respectively; for the independent branches to be trained unique to each model to be trained, the weight parameters of the independent branches to be trained are trained.
[0092] As Figure 5 shown, the electronic device includes two models to be trained. The two models to be trained include the shared branch 1 to be trained, the shared branch 2 to be trained, and the shared branch 3 to be trained. One of the models to be trained includes the unique independent branch 1 to be trained, and the other model to be trained includes the unique independent branch 2 to be trained. The independent branch 1 to be trained and the independent branch 2 to be trained are deleted, and the shared branch 1 to be trained, the shared branch 2 to be trained, and the shared branch 3 to be trained in at least two models to be trained are trained.
[0093] As Figure 6 shown, the parameters of the fixed target sharing branch are trained for the unique independent branch 1 to be trained and the independent branch 2 to be trained respectively, and the target independent branch 1 and the target independent branch 2 are obtained; based on the target sharing branch 1, the target sharing branch 2, the target sharing branch 3, and the target independent branch 1, one trained target model is obtained, and based on the target sharing branch 1, the target sharing branch 2, the target sharing branch 3, and the target independent branch 2, another trained target model is obtained.
[0094] Optionally, the training accuracy of the independent branch to be trained is higher than that of the shared branch to be trained.
[0095] It can be understood that the independent branch to be trained has not been trained when the shared branch to be trained is being trained. Therefore, it only needs to be trained until it converges basically, while the training process of the independent branch to be trained ensures the accuracy and effect of each model, enables the corresponding task to be fully trained, and the loss function is small enough. At the same time, the shared branch acts on multiple subtasks, while the independent branch is for a single task. Therefore, the training accuracy of the independent branch to be trained will be higher, that is, the training accuracy of the independent branch to be trained is higher than that of the shared branch to be trained.
[0096] In this embodiment, the electronic device trains the shared branches to be trained shared by at least two models to be trained to obtain the trained target shared branches, and then fixes the parameters of the target shared branches and trains the independent branches to be trained unique to each model to be trained, which can effectively utilize the parameters of the target shared branches and effectively improve the model accuracy through the independent branches. It can be understood that since the target shared branches need to be shared by multiple target models, the training of the target shared branches needs to be carried out first.
[0097] In one embodiment, a model processing method is further provided, which is applied to an electronic device. The method includes the following steps:
[0098] Step A1, train the shared branch to be trained that is common to at least two models to be trained, and obtain the trained target shared branch.
[0099] Step A2, fix the parameters of the target shared branch, and for each independent branch to be trained that is unique to each model to be trained, train the independent branch to be trained to obtain the target independent branch; the training accuracy of the independent branch to be trained is higher than that of the shared branch to be trained.
[0100] Step A3, based on the target shared branch and the target independent branches, obtain at least two trained target models.
[0101] Step A4, load the target shared branch into the memory of the electronic device; the number of parameters of the target shared branch is less than that of any one of the at least two target models.
[0102] Step A5, determine the target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory.
[0103] Step A6, run the target shared branch and the target independent branch of the target model to be run in the memory, delete the target independent branch in the memory, and return to execute Step A5 until the target shared branch and the target independent branches of the at least two target models are all run, so as to implement the target task.
[0104] In one embodiment, as Figure 7 shown, a model training method is provided. In this embodiment, it is exemplified that this method is applied to an electronic device. The electronic device may be a terminal or a server; it can be understood that this method can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal may be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the electronic device includes at least two models to be trained. The at least two models to be trained include a shared branch to be trained that is common, and an independent branch to be trained that is unique to each model to be trained. The model training method includes the following steps:
[0105] Step S702: Train the shared branch to be trained to obtain the trained target shared branch.
[0106] Optionally, the electronic device deletes the independent branches to be trained in at least two models to be trained to obtain the shared branch to be trained shared by at least two models to be trained; uses the loss function and labels of multi-task to train the shared branch to be trained shared by at least two models to be trained to obtain the trained target shared branch.
[0107] Step S704: Fix the parameters of the target shared branch, and for each independent branch to be trained, train the independent branch to be trained to obtain the target independent branch.
[0108] The electronic device fixes the parameters of the target shared branch, and trains the corresponding independent branches to be trained for different tasks respectively; for the independent branches to be trained unique to each model to be trained, train the weight parameters of the independent branches to be trained.
[0109] Optionally, the training accuracy of the independent branch to be trained is higher than that of the shared branch to be trained.
[0110] Step S706: Based on the target shared branch and the target independent branches, obtain at least two trained target models.
[0111] The electronic device combines the target shared branch and each target independent branch into a target model to obtain at least two trained target models.
[0112] It can be understood that the independent branch to be trained has not been trained when the shared branch to be trained is being trained. Therefore, it is only necessary to train until it converges basically. The training process of the independent branch to be trained ensures the accuracy and effect of each model, enables the corresponding task to be fully trained, and the loss function is small enough. At the same time, the shared branch acts on multiple subtasks, while the independent branch is for a single task. Therefore, the training accuracy of the independent branch to be trained will be higher, that is, the training accuracy of the independent branch to be trained is higher than that of the shared branch to be trained.
[0113] In this embodiment, the electronic device trains the shared branch to be trained shared by at least two models to be trained to obtain the trained target shared branch, then fixes the parameters of the target shared branch, and trains the independent branches to be trained unique to each model to be trained, which can effectively utilize the parameters of the target shared branch and effectively improve the model accuracy through the independent branches.
[0114] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0115] Based on the same inventive concept, an embodiment of the present application also provides a model processing device for implementing the above-mentioned model processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model processing device provided below can refer to the limitations on the model processing method in the above text, and will not be repeated here.
[0116] In an exemplary embodiment, as Figure 8 shown, a model processing device is provided, which is applied to an electronic device. The electronic device includes at least two target models. The at least two target models include a common target shared branch and a target independent branch unique to each target model. The device includes: a target shared branch loading module 802, an independent shared branch loading module 804, and a running module 806, where:
[0117] The target shared branch loading module 802 is configured to load the target shared branch into the memory of the electronic device.
[0118] The independent shared branch loading module 804 is configured to determine the target model to be run from at least two target models, and load the target independent branch of the target model to be run into the memory.
[0119] The running module 806 is configured to run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from at least two target models until the target shared branch and the target independent branch of at least two target models are run out to achieve the target task.
[0120] The above-mentioned model processing device, the electronic device includes at least two target models. The at least two target models include a common target shared branch and a target independent branch unique to each model. The electronic device loads the target shared branch into the memory of the electronic device, determines the target model to be run from the at least two target models, and loads the target independent branch of the target model to be run into the memory. Then, the target shared branch and the target independent branch of the target model to be run can be run in the memory to implement the task corresponding to one of the target models, and the step of determining the target model to be run from the at least two target models is returned until the target shared branch and the target independent branches of the at least two target models are all run. That is to say, during the running of the at least two target models, the target shared branch is always loaded in the memory, and only the target independent branch of the target model to be run needs to be loaded into the memory, so that the target model can be run more quickly, thereby improving the running efficiency of the at least two target models and more quickly achieving the target task.
[0121] In one embodiment, the above-mentioned device further includes a deletion module, and the deletion module is used to delete the target independent branch in the memory.
[0122] In one embodiment, the number of parameters of the target shared branch is less than the number of parameters of any one of the at least two target models.
[0123] In one embodiment, the above-mentioned device further includes a shared branch training module, an independent branch training module, and a target model obtaining module; the shared branch training module is used to train the to-be-trained shared branch common to at least two to-be-trained models to obtain the trained target shared branch; the independent branch training module is used to fix the parameters of the target shared branch, and for the to-be-trained independent branch unique to each to-be-trained model, train the to-be-trained independent branch to obtain the target independent branch; the target model obtaining module is used to obtain at least two trained target models based on the target shared branch and the target independent branch.
[0124] In one embodiment, the training accuracy of the to-be-trained independent branch is higher than the training accuracy of the to-be-trained shared branch.
[0125] Based on the same inventive concept, an embodiment of the present application also provides a model training device for implementing the above-mentioned model training method. The implementation solution provided by this device to solve the problem is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the model training device provided below can refer to the limitations on the model training method in the above text, and will not be repeated here.
[0126] In an exemplary embodiment, as Figure 9As shown, a model training device is provided, which is applied to an electronic device. The electronic device includes at least two models to be trained. The at least two models to be trained include a common shared branch to be trained and an independent branch to be trained unique to each model to be trained. The device includes: a shared branch training module 902, an independent branch training module 904, and a target model obtaining module 906, where:
[0127] The shared branch training module 902 is configured to train the shared branch to be trained to obtain a trained target shared branch.
[0128] The independent branch training module 904 is configured to fix the parameters of the target shared branch and, for each independent branch to be trained, train the independent branch to be trained to obtain a target independent branch.
[0129] The target model obtaining module 906 is configured to obtain at least two trained target models based on the target shared branch and the target independent branches.
[0130] In the above model training device, the electronic device trains the shared branch to be trained that is common to at least two models to be trained to obtain a trained target shared branch, and then fixes the parameters of the target shared branch and trains the independent branch to be trained unique to each model to be trained, which can effectively utilize the parameters of the target shared branch and effectively improve the model accuracy through the independent branches.
[0131] Each module in the above model processing device and model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the electronic device in hardware form or independent of the processor, or stored in the memory in the electronic device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0132] In an exemplary embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as Figure 10As shown in the figure. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a model processing method or a model training method. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0133] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0134] In one embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0135] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0136] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0139] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0140] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A model processing method, characterized in that, Applied to an electronic device, the electronic device includes at least two target models, the at least two target models include a common target shared branch and a target independent branch unique to each model, and the method includes: Load the target shared branch into the memory of the electronic device; Determine a target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory; Run the target shared branch and the target independent branch of the target model to be run in the memory, and return to the step of determining the target model to be run from the at least two target models until the target shared branch and the target independent branch of the at least two target models are run to achieve the target task.
2. The method according to claim 1, wherein After running the target shared branch and the target independent branch of the target model to be run in the memory, it further includes: Delete the target independent branch in the memory.
3. The method according to any one of claims 1 to 2, characterized in that, The number of parameters of the target shared branch is less than the number of parameters of any one of the at least two target models.
4. The method according to claim 1, characterized in that Before loading the target shared branch into the memory of the electronic device, it further includes: Train a to-be-trained shared branch common to at least two to-be-trained models to obtain a trained target shared branch; Fix the parameters of the target shared branch, and train the to-be-trained independent branch for each to-be-trained model unique to obtain a target independent branch; Based on the target shared branch and the target independent branch, obtain at least two trained target models.
5. The method according to claim 4, characterized in that, The training accuracy of the to-be-trained independent branch is higher than the training accuracy of the to-be-trained shared branch.
6. A model training method, characterized in that, Applied to an electronic device, the electronic device includes at least two to-be-trained models, the at least two to-be-trained models include a common to-be-trained shared branch and a to-be-trained independent branch unique to each to-be-trained model, and the method includes: Train the to-be-trained shared branch to obtain a trained target shared branch; Fix the parameters of the target shared branch, and train the to-be-trained independent branch for each to-be-trained independent branch to obtain a target independent branch; Based on the target shared branch and the target independent branch, obtain at least two trained target models.
7. A model processing device, characterized in that, Applied to an electronic device, the electronic device includes at least two target models, the at least two target models include a common target shared branch and a target independent branch unique to each target model, and the device includes: A target shared branch loading module, configured to load the target shared branch into the memory of the electronic device; An independent shared branch loading module, configured to determine a target model to be run from the at least two target models, and load the target independent branch of the target model to be run into the memory; A running module, configured to run the target shared branch and the target independent branch of the target model to be run in the memory, and return to execute the step of determining the target model to be run from the at least two target models until the target shared branches and target independent branches of the at least two target models are run to achieve a target task.
8. A model training device, characterized in that, Applied to an electronic device, the electronic device includes at least two models to be trained, the at least two models to be trained include a common shared branch to be trained and a unique independent branch to be trained for each model to be trained, and the apparatus includes: A shared branch training module, configured to train the shared branch to be trained to obtain a trained target shared branch; An independent branch training module, configured to fix the parameters of the target shared branch, and for each independent branch to be trained, train the independent branch to be trained to obtain a target independent branch; A target model obtaining module, configured to obtain at least two trained target models based on the target shared branch and the target independent branch.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.