A method for improving the speed of running a machine learning model on an embedded device

By dividing the machine learning model into sub-running models and adaptively selecting the model based on the time-consuming list, the problem of low computing efficiency in embedded devices is solved, and the model speed improvement without human intervention is achieved.

CN114519432BActive Publication Date: 2025-07-11ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111605385.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-25
Publication Date
2025-07-11
Estimated Expiration
2041-12-25

AI Technical Summary

Technical Problem

The prior art has low computational efficiency and time-consuming when running machine learning models in embedded devices, and the existing optimization solutions require manual intervention and have poor results.

Method used

Divide the machine learning model into multiple sub-running models, obtain the time-consuming operation of each sub-running model in an embedded device, adaptively select the current running model based on the number of current pending data and the time-consuming list, and send it to the embedded device to run.

Benefits of technology

Improves the speed of running machine learning models on embedded devices, with flexibility and universality, and does not require manual modification of model parameters or structures, and is suitable for any machine learning model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519432B_ABST
    Figure CN114519432B_ABST
Patent Text Reader

Abstract

The present application discloses a method for improving the speed of running a machine learning model on an embedded device. The method includes: obtaining a trained machine learning model; generating a corresponding running model for running on the embedded device based on the machine learning model, where the running model includes a plurality of sub-running models; obtaining the running time consumed by each sub-running model when running on the embedded device to generate a running time list; formulating a current running model based on the quantity of currently to-be-processed data and the running time list, and sending the current running model to the embedded device so that the embedded device runs the current running model, where the current running model includes sub-running models and the function of the current running model is the same as that of the machine learning model. By the above method, the present application can improve the speed of running a machine learning model on an embedded device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and specifically relates to a method for improving the speed of running a machine learning model on an embedded device. Background Art

[0002] With the rapid development of deep learning technology, deep learning technology has been widely applied in fields such as machine vision, artificial intelligence, and embedded devices. However, when deep learning technology is applied to embedded devices, there are still problems such as low computational efficiency and long running time. Therefore, it is necessary to optimize the performance of embedded devices and improve the running speed of embedded devices. The current optimization solutions can only be achieved by modifying network parameters / structures of the convolutional network in deep learning technology, which requires manual implementation, consumes manpower and material resources, and cannot guarantee the optimization effect of the running speed of embedded devices. Summary of the Invention

[0003] This application provides a method for improving the speed of running a machine learning model on an embedded device, which can improve the speed of running a machine learning model on an embedded device.

[0004] To solve the above technical problems, the technical solution adopted by this application is: providing a method for improving the speed of running a machine learning model on an embedded device, the method includes: obtaining a trained machine learning model; generating a corresponding running model for running on the embedded device based on the machine learning model, the running model includes multiple sub-running models; obtaining the running time consumption of each sub-running model running on the embedded device, and generating a running time consumption list; formulating a current running model based on the quantity of the current data to be processed and the running time consumption list, and sending the current running model to the embedded device so that the embedded device runs the current running model, the current running model includes sub-running models, and the function of the current running model is the same as that of the machine learning model.

[0005] To solve the above technical problems, another technical solution adopted by this application is: providing an optimization device, the optimization device includes a memory and a processor connected to each other, wherein, the memory is used for storing a computer program, and when the computer program is executed by the processor, it is used to implement the method for improving the speed of running a machine learning model on an embedded device in the above technical solution.

[0006] To solve the above technical problems, another technical solution adopted by this application is: providing an optimization system, which includes an optimization device and an embedded device connected to each other, the optimization device is the optimization device in the above technical solution, and the embedded device is used for running the current running model.

[0007] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the method for improving the speed of running a machine learning model on an embedded device in the above technical solution.

[0008] Through the above solution, the beneficial effect of this application is as follows: first, obtain the trained machine learning model; then process the machine learning model to generate a running model prepared to run on the embedded device; then obtain the running time consumed by each sub-running model in the running model on the embedded device to generate a running time list; finally, formulate the current running model according to the quantity of the currently to-be-processed data and the running time list, and send the current running model to the embedded device so that the embedded device runs the current running model; since the entire machine learning model is divided into multiple sub-running models and the current running model is adaptively generated according to the running time list of the sub-running models applied in the embedded device, it is possible to select the current running model with less time cost to implement the functions of the machine learning model, improve the speed of running the machine learning model on the embedded device, without the need for manual modification of the parameters / structure of the machine learning model, and at the same time, it is not limited by the type of the machine learning model, having higher flexibility and universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical model in the embodiments of this application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. Among them:

[0010] Figure 1 is a schematic flowchart of an embodiment of the method for improving the speed of running a machine learning model on an embedded device provided by this application;

[0011] Figure 2 is a schematic flowchart of another embodiment of the method for improving the speed of running a machine learning model on an embedded device provided by this application;

[0012] Figure 3 is a schematic flowchart of the process for generating the final running time list provided by this application;

[0013] Figure 4 is a schematic diagram of generating the current running model provided by this application;

[0014] Figure 5 is a schematic structural diagram of an embodiment of the optimization device provided by this application;

[0015] Figure 6 It is a schematic structural diagram of an embodiment of the optimization system provided by this application;

[0016] Figure 7 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners

[0017] The following will further describe this application in detail in conjunction with the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate this application, but do not limit the scope of this application. Similarly, the following embodiments are only partial embodiments of this application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0018] When "embodiment" is mentioned in this application, it means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0019] It should be noted that the terms "first", "second", and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and explicitly defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0020] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the method for improving the speed of running a machine learning model on an embedded device provided by this application. The method includes:

[0021] Step 11: Obtain a trained machine learning model.

[0022] The machine learning model may include a neural network, a support vector machine, a decision tree, a clustering algorithm, etc., which are not limited herein and can be selected and applied according to the actual situation. Obtaining a trained machine learning model can enable the machine learning model to achieve the expected effect when applied to an embedded device, and prevent operation problems from occurring during the application process due to the defects of the machine learning model itself. To achieve a better operation effect, a mature model with a reasonable and effective structure and after pruning and compression processing can also be used as the machine learning model.

[0023] In a specific embodiment, the data to be processed can be obtained first, and then the trained machine learning model can be obtained, so as to perform image processing on the data to be processed through the obtained machine learning model. Specifically, the data to be processed may include images or videos. By using the machine learning model to process the images or videos, tasks such as object detection, object recognition, reconstruction, or tracking can be achieved. It can be understood that in other embodiments, the data to be processed and the training sample data can also be obtained first, and then the machine learning model can be obtained, so as to train the obtained machine learning model through the training sample data to obtain a trained machine learning model, and then perform image processing on the data to be processed through the trained machine learning model. Among them, the training sample data may include image samples or video samples for model training.

[0024] Step 12: Generate a corresponding running model for running in the embedded device based on the machine learning model.

[0025] The machine learning model can be applied to an embedded device to implement corresponding functions. Specifically, when applying the machine learning model to the embedded device, the machine learning model can be first converted into a platform model (i.e., a running model) corresponding to the embedded device through a corresponding conversion chip in the embedded device. The running model can be used as an entity model of the machine learning model in the embedded device to implement the same functions as the machine learning model. For example: If the machine learning model is a classification model, it can classify input data such as images, videos, or voices and output corresponding classification results. When applying this classification model to the embedded device, the embedded device can implement the corresponding classification function by running the corresponding running model.

[0026] Furthermore, the running model may include multiple sub-running models. For example: If the dimension of the trained machine learning model is x, then when converting this machine learning model into a running model, x sub-running models can be generated correspondingly, which are represented by b1, b2,..., bx respectively.

[0027] Step 13: Obtain the running time of each sub-running model running in the embedded device and generate a running time list.

[0028] Each sub - running model can be run on the embedded device to obtain the running time of each sub - running model on the embedded device, thereby generating a running time list. The running time list can include all sub - running models and their respective running times, as shown in the following table:

[0029]

[0030] Step 14: Develop the current running model based on the quantity of the current data to be processed and the running time list, and send the current running model to the embedded device so that the embedded device runs the current running model.

[0031] Due to the influence of factors such as the application scenario, in the actual application process, the quantity of the data to be processed input to the running model will change. Developing the current running model according to the quantity of the current data to be processed and the running time list can select the running model with the minimum running time in the current application scenario from multiple sub - running models as the current running model (the function of this current running model is the same as that of the machine - learning model), and send the current running model to the embedded device so that the embedded device runs the current running model, thereby improving the running speed of the embedded device.

[0032] Furthermore, the current running model includes sub - running models. The number of sub - running models included in the current running model can be one, two, three, or more than three. The dimension and quantity of the sub - running models it contains can be selected according to the quantity of the current data to be processed and the running time. It can be understood that the number of running times corresponding to each sub - running model included in the current running model can also be reasonably allocated according to the current running model and the quantity of the data to be processed. For example: the current sub - running models include b1 - b8, and the quantity of the current data to be processed is 10. If it is determined that the current running model is b1, then the sub - running model b1 can be run ten times. If it is determined that the current running model is b1 and b8, then the sub - running model b1 can be run twice and the sub - running model b8 can be run once respectively.

[0033] In the solution provided in this embodiment, first, the running model corresponding to the machine - learning model deployed to the embedded device is counted; then, the running time of each sub - running model in the running model is tested to obtain a running time list; then, according to the quantity of the data to be processed in the current application environment and the running time list, the current running model with the shortest running time is adaptively developed. Using this current running model to execute the same task as the machine - learning model can improve the running speed of the embedded device running the machine - learning model, without the need for manual modification of the parameters / structure of the machine - learning model, has higher flexibility, and is not restricted by the type of the machine - learning model. It can be applied to any application scenario of the machine - learning model and has high universality.

[0034] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the method provided in this application for improving the speed of running a machine learning model on an embedded device. The method includes:

[0035] Step 21: Obtain a trained machine learning model.

[0036] The above step 21 is the same as step 11 in the above embodiment and will not be elaborated here.

[0037] Step 22: Generate a corresponding running model for running on the embedded device based on the machine learning model.

[0038] The data input into the machine learning model has a size, which can be represented by N i C i H i W i where N i represents the number of input data, C i represents the channels of the input data, H i represents the height of the input data, and W i represents the width of the input data. By processing the input data using the machine learning model, result data with a size of N0C0H0W0 can be output. The ratio of the input data size N i C i H i W i of the machine learning model to the output data size N0C0H0W0 is fixed. It can be understood that the ratio of the input data size to the output data size during the running of the running model is the same as that of the machine learning model. Taking the number of data N as an example, if the number of input data N i of the machine learning model is 1 and the number of output data N0 is 3, then when the number of input data of the running model is 2, the number of its output data will be 6.

[0039] Furthermore, the running model may include multiple sub-running models. The number of sub-running models is the same as the dimension of the machine learning model. The maximum value of the dimensions of the multiple sub-running models is equal to the number of machine learning models. The dimension of the machine learning model is the number of data that the machine learning model can process simultaneously (i.e., N of the above data size). The dimension of the trained machine learning model is a fixed value. For example, taking the input data as pictures, if the dimension of the trained machine learning model is x, then when converting this machine learning model into a running model, x sub-running models can be correspondingly generated, represented by b1, b2, …, bx respectively. The dimensions of the x sub-running models b1, b2, …, bx are 1, 2, …, x respectively. That is, the sub-running model b1 can process one picture simultaneously, b2 can process two pictures simultaneously, and so on, bx can process x pictures simultaneously.

[0040] Step 23: Obtain the running time consumption of each sub-running model when running in the embedded device, and generate a running time consumption list.

[0041] Each sub-running model can be run in the embedded device for a running test to obtain the running time consumption of each sub-running model when running in the embedded device, so as to generate a running time consumption list. Then, based on the number of currently to-be-processed data and the running time consumption list, the current running model is formulated, so that the embedded device runs the current running model. The steps of formulating the current running model based on the number of currently to-be-processed data and the running time consumption list are introduced below, which specifically include steps 24 to 26:

[0042] Step 24: Screen out at least one preferred sub-running model from the multiple sub-running models based on the running time consumption list.

[0043] The running time consumption list includes the sub-running models and the corresponding running time consumption. Based on the running time consumption list, multiple sub-running models generated by converting the embedded device can be deleted. The sub-running models with longer running time are deleted, and at least one preferred sub-running model with shorter running time is screened out. At the same time, the information of the corresponding sub-running models in the running time consumption list is updated to generate a final running time consumption list containing at least one preferred sub-running model, so that after the embedded device receives the final running time consumption list, it runs the corresponding preferred sub-running model, thereby achieving the improvement of the running speed of the embedded device.

[0044] In a specific embodiment, a sub - running model can be first selected from multiple sub - running models as the candidate model; then the candidate model is compared with other sub - running models in the multiple sub - running models in turn to obtain a comparison result; then a preferred sub - running model is generated based on the comparison result; specifically, the comparison result may include a list of final running times, and the preferred sub - running model is the sub - running model included in the list of final running times. Specifically, the candidate model can be first compared with other sub - running models in turn for a first screening to obtain an intermediate running time list; then the sub - running models in the intermediate running time list are screened to obtain the final running time list, that is, the final running time list is generated by screening the running time list twice.

[0045] Further, when the candidate model is compared with other sub - running models in turn, if the dimension of the candidate model is smaller than the dimension of other sub - running models, and the running time of the candidate model is greater than the running time of other sub - running models, then the candidate model is deleted from the running time list, and the step of selecting a sub - running model from multiple sub - running models as the candidate model is returned until all the multiple sub - running models are traversed to generate an intermediate running time list.

[0046] For example: the current sub - running models include b1 to b10, and their corresponding running times are T1 to T10 respectively. b1 can be used as the candidate model and compared with other sub - running models b2 to b10 in turn. Then, in the comparison process, the situation of T1>T3 occurs, that is, the time taken to run the sub - running model b1 to process one data is longer than the time taken to run the sub - running model b3 to process three data at the same time. At this time, the sub - running model b1 with a longer running time is deleted from the running time list, and then the next sub - running model b2 is used as the candidate model and compared with other sub - running models, and so on, until all the sub - running models in the running time list are polled to obtain an updated intermediate running time list.

[0047] The above - mentioned first screening of multiple sub - running models in the running time list is carried out under the condition of processing different amounts of data, and only the situation where the dimension of the candidate model is smaller than the dimension of other sub - running models and the running time of the candidate model is greater than the running time of other sub - running models can be avoided. However, among the sub - running models in the intermediate running time list generated after the first screening, there may still be a situation where the running time of a sub - running model with a smaller dimension is smaller than the running time of a sub - running model with a larger dimension, but when processing the same amount of data, there may still be a situation where the running time required to run the sub - running model with a smaller dimension multiple times is smaller than the running time required to run the sub - running model with a larger dimension once. Therefore, at this time, the sub - running models in the intermediate running time list are screened a second time to delete the redundant sub - running models with a larger dimension in the above - mentioned situation. The specific steps are as followsFigure 3 As shown in:

[0048] Step 31: Select a sub-run model in the intermediate running time-consuming list as the candidate model.

[0049] Step 32: Calculate the ratio of the dimension of the remaining model to the dimension of the candidate model to obtain the dimension ratio.

[0050] The remaining models are the sub-run models in the intermediate running time-consuming list except the candidate model. Taking the intermediate running time-consuming list containing sub-run models b1 and b8 as an example, calculate the ratio of the dimension of the candidate model to the dimension of the remaining model. b1 can be used as the candidate model, compare it with the remaining model b8, calculate the ratio of the dimension of b8 to b1, and obtain the dimension ratio of 8. Then, based on the dimension ratio, calculate the running time required for the candidate model when processing the same number of data.

[0051] Step 33: Multiply the running time of the candidate model by the dimension ratio to obtain the product value.

[0052] Taking the dimension ratio of the above remaining model b8 to the candidate model b1 as 8 as an example, at this time, multiply the running time T1 of the candidate model b1 by the dimension ratio 8, and the running time 8*T1 required for the candidate model b1 when processing eight data can be obtained, that is, the product value. Thus, under the condition of processing the same number of data, compare it with the running time T8 of b8.

[0053] Specifically, when calculating the dimension ratio of the candidate model and the remaining model, it can be judged whether the dimension ratio is greater than the preset value. Among them, the preset value can be 1, that is, judge whether the dimension ratio is greater than 1. If the dimension ratio is greater than 1, then execute the step of judging whether the product value is less than the running time of the corresponding sub-run model; if the dimension ratio is less than 1, then return to the step of selecting a sub-run model in the intermediate running time-consuming list as the candidate model to re-select a new candidate model; it can be understood that since the running time of the sub-run model refers to the time required to run a sub-run model once, the running time of each sub-run model is a fixed value, and it will not become smaller because the number of data to be processed becomes smaller. That is, although a sub-run model with a dimension of 3 is used to process one data, the time required for it is the same as the time required for it to process three data. Therefore, when calculating the running time required for each of them when processing the same number of data based on the dimension ratio and then comparing the running times, the data processing ability of the sub-run model with a larger dimension should be used as the benchmark for comparing the running times. That is, when calculating the product data for the candidate model, the dimension ratio used is greater than 1; for example: when comparing b1 and b8, the running times required for both of them to process eight data can be compared.

[0054] Step 34: Judge whether the product value is less than the running time of the remaining model.

[0055] After calculating the product value, that is, calculating the running time required for the candidate model to process the same amount of data as the remaining models, comparing it with the running time of the remaining models, and determining whether the product value is less than the running time of the remaining models, so as to determine whether there is a situation where the running time required to run a sub-running model with a larger dimension once is greater than the running time required to run multiple sub-running models with smaller dimensions.

[0056] Step 35: If the product value is less than the running time of the remaining models, delete the remaining models from the intermediate running time list.

[0057] When the product value is less than the running time of the remaining models, delete the remaining models with larger dimensions from the intermediate running time list, retain the candidate model with less running time, and return to the step of selecting a sub-running model as the candidate model from the intermediate running time list until the intermediate running time list is traversed, and generate a final running time list containing at least one preferred sub-running model; for example: the candidate model is b1, its running time is T1, the remaining model is b4, its running time is T4, at this time the calculated dimension ratio of the remaining model to the candidate model is 4, and the value is greater than 1, then multiply the running time T1 of the candidate model by the dimension ratio 4 to get the product value 4*T1, which means the time required for the candidate model to process four data, and then compare the product value 4*T1 with the running time T4 of the remaining model. If 4*T1 is less than T4, it means that the time required to run b1 four times respectively is shorter than the time required to run b4 once, then the remaining model b4 can be deleted at this time, and the b1 with less running time can be retained.

[0058] Step 25: Combine at least one preferred sub-running model based on the current number of data to be processed to generate at least one candidate running model.

[0059] The candidate running model includes at least one preferred sub-running model and the running times of the preferred sub-running models. The sum of the products of the dimensions of all preferred sub-running models and the corresponding running times is greater than or equal to the current number of data to be processed; specifically, taking the preferred sub-running models b1 and b4 included in the final running time list and the current number of data to be processed as 10 as an example, running the current preferred sub-running models to process 10 data, the candidate running models can include the following running schemes: running b1 ten times, running b1 six times and b4 once, running b4 twice and b1 twice, running b4 three times, etc.

[0060] Understandably, the data vacancies in b4 can be filled by adding blank data, so that data with a quantity less than four can be processed using b4. At this time, a scheme similar to running b1 eight times and b4 once or running b4 three times can also be adopted to process ten currently to-be-processed data.

[0061] Step 26: Select the candidate running model with the minimum total running time among all candidate running models as the current running model.

[0062] Calculate the sum of the products of the running time of all preferred sub-running models included in the candidate running model and the corresponding number of runs to obtain the total running time of the candidate running model. Select the one with the minimum total running time from all candidate running models as the current running model. For example: There are a first candidate running model and a second candidate running model among the candidate running models. The first candidate running model includes a preferred sub-running model b1 (with a running time of T1), and the corresponding number of runs is 10. Then the total running time of the first candidate running model is 10*T1. The second candidate running model includes a preferred sub-running model b1 (with a running time of T1) and b4 (with a running time of T4), and the corresponding number of runs are 6 and 1 respectively. Then the total running time of the second candidate running model is 6*T1 + T4. If (6*T1 + T4) is less than 10*T1, it means that the total running time of the second candidate running model is smaller, and the second candidate running model can be used as the current running model.

[0063] For example, as Figure 4 shown, taking the number of to-be-processed data as 9 as an example, the preferred sub-running models include b1 to b8, and their respective corresponding running times are T1 to T8. Combine multiple preferred sub-running models according to the number of currently to-be-processed data to generate multiple candidate running models. For example Figure 4 in: "b1 + b8", "b2 + b7", "b3 + b6", and "b4 + b5", so as to complete the processing of these 9 to-be-processed data. The total running time corresponding to each candidate running model is Ta (Ta = T1 + T8), Tb (Tb = T2 + T7), Tc (Tc = T3 + T6), and Td (Td = T4 + T5) respectively. Then select the candidate running model with the minimum total running time among Ta to Td as the current running model. Understandably, Figure 4 only the above four candidate running models are used as examples for illustration. In the actual application scenario, the candidate running model can also include various combination methods such as 5*b1 + b4, and then select the candidate running model with the minimum total running time among all candidate running models as the current running model.

[0064] Step 27: Package all the preferred sub-running models included in the current running model to generate compressed data, and send the compressed data to the embedded device so that the embedded device runs the preferred sub-running models based on the current running model.

[0065] Package the preferred sub-running models and send them to the embedded device so that the embedded device runs the preferred sub-running models based on the current running model with the least time consumption, thereby improving the running speed of the embedded device applying the machine learning model; it can be understood that the embedded device needs corresponding running components to run the sub-running models. When the number of preferred sub-running models included in the current running model is multiple, the embedded device can select to run multiple preferred sub-running models simultaneously or run multiple preferred sub-running models sequentially according to the number of running components.

[0066] In this embodiment, at least one preferred sub-running model is screened out from multiple sub-running models according to the running time consumption list. By screening the multiple sub-running models twice, the sub-running models with relatively large time consumption in the multiple sub-running models are deleted, and the preferred sub-running models are screened out; different from the related technology of adjusting the running speed of the embedded device by adjusting the size height / width of the data, this solution divides the quantity of the data to be processed, and then selects the candidate running model with the least total running time consumption from all candidate running models as the current running model according to the current quantity of the data to be processed, so that the embedded device runs the current running model, and finally realizes the improvement of the running speed of the embedded device. There is no need for manual intervention to divide the quantity of the data to be processed, and it can achieve fully automated test acquisition, reducing the extra workload introduced by different chips among embedded devices; moreover, by adopting the method of automatically measuring the running time consumption of the sub-running models in the embedded device, using the actual measured time data to represent the hardware performance of the embedded device, avoiding the hardware differences of different embedded devices, the running time consumption list of different embedded devices in their respective application environments can be obtained, thereby realizing the adaptation to all types of embedded devices and providing optimization solutions suitable for their respective performances and application environments for different types of embedded devices; furthermore, there is no need for manual participation in measurement and calculation, greatly reducing the workload required by the embedded device, being able to realize the adaptive adjustment of the models running on the embedded device, saving manpower and material resources; in addition, this embodiment does not need to improve the algorithm of the machine learning model itself, and the implementation is simple.

[0067] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an embodiment of the optimization device provided by this application. The optimization device 50 includes a memory 51 and a processor 52 that are connected to each other. The memory 51 is used to store a computer program, and when the computer program is executed by the processor 52, it is used to implement the method for improving the running speed of the machine learning model on the embedded device in the above embodiment.

[0068] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of an optimization system provided by the present application. The optimization system 60 includes an interconnected optimization device 61 and an embedded device 62. The optimization device 61 is the optimization device in the above technical solution, and the embedded device 62 is used to run the currently running model.

[0069] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 70 is used to store a computer program 71, which, when executed by a processor, is used to implement the method for improving the speed of running a machine learning model on an embedded device in the above embodiment.

[0070] The computer-readable storage medium 70 can be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0071] In several implementation manners provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manner described above is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0072] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the model of this implementation manner.

[0073] In addition, each functional unit in various implementation manners of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0074] The above are only embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for improving the speed of running a machine learning model on an embedded device, characterized in that, Including: Obtain a trained machine learning model, and use the machine learning model to perform image processing or video processing on the image data to be processed; Generate a running model corresponding to the operation in the embedded device based on the machine learning model, where the running model includes a plurality of sub-running models; Obtain the running time consumed by each of the sub-running models when running in the embedded device, and generate a running time list; Formulate a current running model based on the quantity of the current image data or video data to be processed and the running time list, and send the current running model to the embedded device, so that the embedded device runs the current running model. The current running model includes the sub-running models, and the function of the current running model is the same as the function of the machine learning model; The step of formulating the current running model based on the quantity of the current image data or video data to be processed and the running time list includes: Screen out at least one preferred sub-running model from the plurality of sub-running models based on the running time list; Combine the at least one preferred sub-running model according to the quantity of the current image data or video data to be processed, and generate at least one candidate running model; Use the candidate running model with the minimum total running time among all the candidate running models as the current running model.

2. The method for improving the speed of running a machine learning model on an embedded device according to claim 1, wherein The step of screening out at least one preferred sub-running model from the plurality of sub-running models based on the running time list includes: Select a sub-running model from the plurality of sub-running models as a candidate model; Compare the candidate model with other sub-running models in the plurality of sub-running models in sequence to obtain a comparison result; Generate the preferred sub-running model based on the comparison result.

3. The method for improving the speed of running a machine learning model on an embedded device according to claim 2, wherein Each of the sub-running models corresponds to a different dimension, and the method further includes: If the dimension of the candidate model is less than the dimension of the other sub-running models, and the running time consumed by the candidate model is greater than the running time consumed by the other sub-running models, then delete the candidate model from the running time list, and return to the step of selecting a sub-running model from the plurality of sub-running models as a candidate model until all the plurality of sub-running models are traversed, and generate an intermediate running time list.

4. The method for improving the speed of running a machine learning model on an embedded device according to claim 3, wherein The comparison result includes a final running time list, and the preferred sub-running model is the sub-running model included in the final running time list; the method further includes: Select a sub-running model from the intermediate running time list as a candidate model; Calculate the ratio of the dimension of the remaining models to the dimension of the candidate model to obtain a dimension ratio, where the remaining models are the sub-running models other than the candidate model in the intermediate running time list; Multiply the running time consumed by the candidate model by the dimension ratio to obtain a product value; Determine whether the product value is less than the running time consumed by the remaining models; If so, delete the remaining model from the intermediate running time consumption list, and return to the step of selecting one of the sub-running models in the intermediate running time consumption list as the candidate model until the traversal of the intermediate running time consumption list is completed, and generate the final running time consumption list including at least one of the sub-running models.

5. The method for improving the speed of running a machine learning model on an embedded device according to claim 4, wherein The method further includes: Determine whether the dimension ratio is greater than a preset value; If so, execute the step of determining whether the product value is less than the running time consumption of the corresponding sub-running model; If not, return to the step of selecting one of the sub-running models in the intermediate running time consumption list as the candidate model.

6. The method for improving the speed of running a machine learning model on an embedded device according to claim 1, wherein The method further includes: The candidate running model includes at least one of the preferred sub-running models and the number of times the preferred sub-running model runs, and the sum of the products of the dimensions of all the preferred sub-running models and the corresponding number of times is greater than or equal to the quantity of the current image data or video data to be processed.

7. The method for improving the speed of running a machine learning model on an embedded device according to claim 6, wherein Before the step of selecting the candidate running model with the minimum total running time consumption as the current running model from the at least one candidate running model, it includes: Calculate the sum of the products of the running time consumption of all the preferred sub-running models included in the candidate running model and the corresponding number of times to obtain the total running time consumption of the candidate running model.

8. The method for improving the speed of running a machine learning model on an embedded device according to claim 7, characterized in that, The method further includes: Package all the preferred sub-running models included in the current running model to generate compressed data, and send the compressed data to the embedded device so that the embedded device runs the preferred sub-running model based on the current running model.

9. The method for improving the speed of running a machine learning model on an embedded device according to claim 1, wherein The method further includes: The number of the sub-running models is the same as the dimension of the machine learning model, and the maximum value of the dimensions of the multiple sub-running models is equal to the number of the machine learning models.

10. An optimization device, characterized in that, It includes a memory and a processor connected to each other. The memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the method for improving the running speed of the machine learning model on the embedded device according to any one of claims 1-9.

11. An optimization system, characterized in that, It includes an optimization device and an embedded device connected to each other. The optimization device is the optimization device described in claim 10 above, and the embedded device is used to run the current running model.

12. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the method for improving the running speed of the machine learning model on the embedded device according to any one of claims 1-9.

Citation Information

Patent Citations

  • Learned resource consumption model for optimizing big data queries

    CN113711198A

  • Method and system for predicting patient outcomes using multi-modal input with missing data modalities

    US20190325995A1