A Client-Based AI Model Dynamic Optimization and Adaptive Inference Method
By dynamically optimizing the AI model and selecting the optimal inference path on the user-side device, the problems of resource differences and load changes in different devices are solved, efficient AI inference performance and efficiency under limited resource conditions are achieved, and the flexibility of AI model deployment is improved.
Patent Information
- Application Number
- CN202510406125.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing technology introduces multi-dimensional parallel processing in the AI model training and inference process to reduce the consumption of AI's computing resources, but it fails to effectively solve the situation where the hardware resources of different user-side devices are large and the changes in device load, resulting in the AI model being unable to be dynamically updated and optimized, and cannot adapt to the resource limitations of client devices, which in turn affects processing speed and effect.
A dynamic optimization and adaptive inference method based on client-side AI model is proposed. By analyzing the hardware resources and operating environment of the user-side device, selecting the appropriate AI model to be deployed to the user-side device, and monitoring the operating status of the device in real time, dynamically optimizing the AI model to select the optimal inference path.
Under limited resource conditions, the user-side equipment adjusts the optimal inference path in real time to ensure the efficient inference performance and efficiency of the AI model. By matching tasks to the cloud, it reduces device workload and resource waste, and improves the flexibility and perfection of the AI model deployment.
Smart Images

Figure CN119903928B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model optimization, and specifically relates to a method for dynamically optimizing an AI model and adaptive inference based on a client. Background Art
[0002] With the development of mobile devices and edge computing, AI models are gradually being transferred from the cloud to client devices for processing. However, client devices have limitations in terms of computing resources, memory, and battery life. How to efficiently run complex AI models under these constraints has become a challenge. Therefore, this application proposes a method for dynamically optimizing an AI model and adaptive inference based on a client.
[0003] The prior art, such as an invention patent application with publication number CN114035936A, discloses a multi-dimensional parallel processing system and method based on artificial intelligence. Through data parallelism, it automatically manages the data to be processed and distributes the data to be processed to hardware processors; sequence parallelism, divides and distributes the data, and places each data to be processed on multiple processors; pipeline parallelism, divides the model into multiple segments, deploys each segment on different hardware processors, and connects them in series according to the model order, and multi-dimensional model parallelism, performs network model partitioning on the training model of the data to be processed scheduled to the processor, schedules the training model to multiple processors, and the optimizer updates the parameters of the model to complete the training process. During the inference process, the above resource scheduling and multi-dimensional parallel technologies are also adopted. By introducing multi-dimensional parallel processing during the training and inference of the AI model, the consumption of computing resources by AI is reduced, the deployment efficiency of artificial intelligence is improved, and the deployment cost is minimized.
[0004] Regarding the above solution, there are the following technical problems: 1. The current technology mainly introduces multi-dimensional parallel processing during the training and inference of the AI model to reduce the consumption of computing resources by AI, so as to improve the deployment efficiency of artificial intelligence and minimize the deployment cost. It does not consider that when the hardware resources of different client devices vary greatly, and thus it is unable to support multiple hardware processors to deploy the model. Nor does it consider the problem of dynamically updating the model according to the load and resource changes of the device during the operation of the client device. The neglect of the above aspects will cause the AI model to be unable to adapt to the resource limitations of the client device, and thus unable to achieve the ideal processing speed and processing effect.
[0005] 2. The current technology lacks real-time monitoring of the running state of the client device during the task processing, and thus lacks the analysis of the optimal inference path of the model based on the real-time monitoring of the running state. The neglect of the above aspects causes the AI model to be unable to automatically adjust according to the hardware resources and running state of the device, and thus unable to ensure the maximization of the inference performance and working efficiency of the AI model. Summary of the Invention
[0006] The purpose of this application is to provide a method for dynamically optimizing and adaptively inferring an AI model based on a client, which solves the problems existing in the background technology.
[0007] To solve the above technical problems, this application adopts the following technical solutions: This application provides a method for dynamically optimizing and adaptively inferring an AI model based on a client, including: Step 1, model selection: Analyze the hardware resource information and operating environment information of the user terminal device, and then select an AI model to be deployed to the user terminal device.
[0008] Step 2, model and inference path optimization: Real-time monitor and analyze the operating state information of the user terminal device, then optimize the AI model, and analyze to obtain an adaptive algorithm to select the optimal inference path.
[0009] Step 3, task matching: Store each executed task of the user terminal device in the cloud. When there is a task to be executed on the user terminal device, match the task to be executed with each executed task in the cloud to obtain the matching task of the task to be executed, and then directly call the AI model and the optimal inference path of the matching task to process the task to be executed.
[0010] Preferably, the hardware resource information of the user terminal device includes the available amount of CPU resources, GPU resources, memory resources, and network status of the user terminal device, and the operating environment information includes the battery power and network bandwidth of the user terminal device.
[0011] Preferably, analyze the resource requirement evaluation coefficients of each AI model based on the running basic information of the AI model, compare the resource requirement evaluation coefficients of each AI model with the available resource evaluation coefficient of the user terminal device, and then obtain the ratio of the resource requirement evaluation coefficient of each AI model to the available resource evaluation coefficient of the user terminal device. Obtain each AI model corresponding to when the ratio is less than 1, record it as each selectable AI model, obtain the inference accuracy of each selectable AI model from the cloud, arrange the inference accuracies of each selectable AI model in descending order, and select the AI model with the highest inference accuracy to be deployed to the user terminal device.
[0012] Preferably, the process of analyzing and obtaining an adaptive algorithm is as follows: A1. Obtain the available amount of CPU resources, GPU resources, memory resources, and network status of the user terminal device based on the hardware resource information of the user terminal device, and obtain the path number, CPU resource occupancy, GPU resource occupancy, memory resource occupancy, and network dependency status of each inference path based on the basic information of each inference path.
[0013] A2. Traverse each inference path. When the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy of a certain inference path are all less than or equal to the available CPU resources, available GPU resources, and available memory resources of the user device, and the network dependency status of this inference path conforms to the network status of the user device, mark this inference path as an available inference path. Accordingly, obtain all available inference paths in the user device, denoted as each available inference path; add up the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy in each available inference path to obtain the total resource occupancy of each available path.
[0014] A3. After normalizing the total resource occupancy and inference accuracy of each available inference path, denote them respectively as and , where is the number corresponding to each available inference path, , is any integer greater than 2. According to the calculation formula: analyze and obtain the priority selection index of each available inference path, where and respectively represent the weight factor corresponding to the total resource occupancy of the inference path and the weight factor corresponding to the inference accuracy.
[0015] Preferably, the process of selecting the optimal inference path is as follows: Sort the priority selection indices of each available inference path in descending order, construct a priority selection table of available inference paths for the AI model, and mark the available inference path corresponding to the maximum priority selection index as the optimal inference path.
[0016] Preferably, the process of matching the task to be executed with each executed task in the cloud to obtain the matching task of the task to be executed is as follows: Obtain the function description set and the resource requirement type set of the task to be executed from the user device; obtain the function description set and the resource requirement type set of each executed task from the cloud, where is the number corresponding to each executed task, , is any integer greater than 2. According to the calculation formula: analyze and obtain the fitness of the task to be executed and each executed task, where and respectively represent the weight factor corresponding to function matching when the user device executes the task and the weight factor corresponding to resource requirement type matching.
[0017] Compare the fitness of the to-be-executed task obtained with the fitness of each executed task with the set task fitness threshold. When the fitness of the to-be-executed task with a certain executed task is greater than or equal to the set task fitness threshold, it is determined that the to-be-executed task and the executed task match successfully; otherwise, it is determined that the to-be-executed task and the executed task do not match successfully. Based on this, obtain each executed task that matches the to-be-executed task successfully, and record the executed task with the highest fitness with the to-be-executed task as the matching task of the to-be-executed task.
[0018] The beneficial effects of this application are as follows: 1. A method for dynamic optimization and adaptive inference of an AI model based on a client provided by this application analyzes the hardware resources and operating environment of the user-side device, and then selects a suitable AI model to be deployed to the user-side device. Then, it monitors the operating status information of the user-side device in real time, so as to dynamically optimize and adjust the AI model to achieve the purpose of efficient inference. At the same time, according to the hardware resource information of the user-side device, the resource occupancy of each inference path, and the inference accuracy, an adaptive algorithm is analyzed. According to this adaptive algorithm, the user-side device can adjust the optimal inference path in real time, so as to ensure the best inference performance and efficiency under limited resource conditions. Finally, each processed task and its corresponding AI model and optimal inference path are stored in the cloud for subsequent direct invocation.
[0019] 2. This application analyzes the hardware resource information and operating environment information of the user-side device, which lays a foundation for the subsequent selection of the AI model. Moreover, the hardware resources of different user-side devices vary greatly, which will lead to great differences in the running speed and effect of the AI model on different client devices. Therefore, analyzing the hardware resource information and operating environment information of the user-side device can select a suitable AI model to be deployed to the user-side device, thereby ensuring the running speed and effect of the AI model.
[0020] 3. This application monitors and analyzes the operating status information of the user-side device after the AI model is deployed, and then optimizes the AI model, so as to adjust the calculation amount and complexity of the AI model in real time according to the operating status information of the user-side device, ensuring that a high-precision large model runs on a user-side device with strong hardware resources, and switching to a lightweight model on a resource-constrained user-side device to ensure the efficiency of inference, which also greatly improves the flexibility of AI model deployment on the user-side device.
[0021] 4. This application analyzes the hardware resources of the user-side device, the resource requirements of each inference path, and the inference accuracy, and then selects the optimal inference path, and adjusts the optimal inference path in real time according to the changes in the hardware resources of the user-side device to ensure the efficiency and stability during the inference process of the AI model.
[0022] 5. This application saves each executed task of the client device to the cloud. When the client device receives a new unexecuted task, it matches the unexecuted task with the executed tasks to obtain the matching tasks of the unexecuted task, and then directly calls the AI model and the optimal inference path of the matching task to process the task to be executed, greatly reducing the workload of the client device and the waste of resources caused by unnecessary analysis, thereby ensuring the perfection and flexibility of this method. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of the implementation steps of the method of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0026] Referring to Figure 1 As shown, this application provides an AI model dynamic optimization and adaptive inference method based on the client, including the following steps: Step 1. Model selection: Analyze the hardware resource information and operating environment information of the client device, and then select an AI model to be deployed to the client device.
[0027] It should be noted that the client device includes smart phones, smart homes, wearable devices, etc.
[0028] In a specific example, the hardware resource information of the client device includes the available amount of CPU resources, GPU resources, memory resources, and network status of the client device, and the operating environment information includes the battery power and network bandwidth of the client device.
[0029] In a specific example, the analysis of the hardware resources and operating environment of the client device is as follows: Based on the hardware resource information of the client device, extract the available amount of CPU resources, GPU resources, and memory resources of the client device, and then analyze to obtain the hardware resource evaluation coefficient of the client device .
[0030] Extract the battery power and network bandwidth of the client device based on the operating environment information of the client device, and then analyze to obtain the operating environment evaluation coefficient of the client device .
[0031] Integrate the hardware resource evaluation coefficient of the client device and the operating environment evaluation coefficient , according to the calculation formula: Analyze to obtain the available resource evaluation coefficient of the client device , where and respectively represent the weight factor corresponding to the hardware resource evaluation coefficient of the client device and the weight factor corresponding to the operating environment evaluation coefficient
[0032] It should be noted that , , .
[0033] It should be noted that the weight factors corresponding to the hardware resource evaluation coefficient of the client device and the weight factor corresponding to the operating environment evaluation coefficient are obtained through factor analysis. First, the information condensation of the spatial path attenuation of the hardware resource evaluation coefficient and the operating environment evaluation coefficient of the client device is performed, and then the variance interpretation rate after rotation is obtained. The weight is obtained by dividing the cumulative variance interpretation rate
[0034] It should be noted that factor analysis is a well-known technology. It is a multivariate statistical analysis method that starts from the study of the internal correlation dependence relationship of variables and reduces a number of variables with intricate relationships into a few comprehensive factors; information condensation is expressed as calculating the median; the variance interpretation rate is the amount of information extracted by the factor, and the variance interpretation rate = eigenvalue / total number of analysis items; the variance interpretation rate after rotation is expressed as the variance interpretation rate of the factor after maximum variance rotation
[0035] In a specific example, the analyzed hardware resource evaluation coefficient of the client device and the operating environment evaluation coefficient , the specific analysis process is as follows: After normalizing the available CPU resources, available GPU resources, and available memory resources of the client device, they are respectively recorded as , and , according to the calculation formula: Analyze to obtain the hardware resource evaluation coefficient of the client device , where , and respectively represent the weight factor corresponding to the available CPU resources, the weight factor corresponding to the available GPU resources, and the weight factor corresponding to the available memory resources; similarly, according to the battery power and network bandwidth of the user device, the operation environment evaluation coefficient of the user device can be analyzed and obtained 。
[0036] It should be noted that , , , ,where 、 and are set in the same way as 、 ,so they will not be elaborated here
[0037] In a specific example, the process of selecting an AI model to be deployed to the user device is as follows: analyze the resource requirement evaluation coefficient of each AI model based on the running basic information of the AI model, compare the resource requirement evaluation coefficient of each AI model with the available resource evaluation coefficient of the user device, and then obtain the ratio of the resource requirement evaluation coefficient of each AI model to the available resource evaluation coefficient of the user device. Obtain each AI model corresponding to a ratio less than 1, record it as each selectable AI model, obtain the inference accuracy of each selectable AI model from the cloud, arrange the inference accuracy of each selectable AI model in descending order, and select the AI model with the largest inference accuracy to be deployed to the user device
[0038] It should be noted that the running basic information of the AI model includes the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy. The analysis method of the resource requirement evaluation coefficient of the AI model is the same as that of the hardware resource evaluation coefficient of the user device, so it will not be elaborated here
[0039] Step 3: Model and inference path optimization: Monitor and analyze the running state of the user device in real time, then optimize the AI model, and analyze and obtain an adaptive algorithm to select the optimal inference path
[0040] In a specific example, the process of monitoring and analyzing the running state information of the user device in real time is as follows: obtain the CPU temperature, CPU load, and battery power consumption per unit time of the user device at the current moment from the monitored running data of the user device, and record them as 、 and , according to the calculation formula: Analyze and obtain the running state evaluation coefficient of the user terminal device at the current moment, where , and represent the standard values of the CPU temperature, CPU load, and power consumption per unit time of the battery at the current moment of the client device, , and represent the weight factors corresponding to the CPU temperature, CPU load, and power consumption per unit time of the battery at the current moment of the client device, respectively.
[0041] Compare the operation state evaluation coefficient of the client device with the threshold of the operation device evaluation coefficient of the client device. When the operation state evaluation coefficient of the client device is greater than or equal to the threshold of the operation state evaluation coefficient of the client device, optimize the AI model; otherwise, do not optimize the AI model.
[0042] It should be noted that the operation data of the client device includes CPU temperature, CPU load, and power consumption per unit time of the battery.
[0043] It should be noted that the standard values of the CPU temperature, CPU load, and power consumption per unit time of the battery at the current moment of the client device are the CPU temperature value, CPU load value, and power consumption per unit time of the battery of the client device before the deployment of the AI model.
[0044] It should be noted that , , , , where , and are set in the same way as , , so it will not be elaborated here.
[0045] It should be noted that the threshold of the operation device evaluation coefficient of the client is the operation evaluation coefficient of the client device obtained by analysis when the available CPU resources, available GPU resources, available memory resources, battery power, or network bandwidth of the client device are insufficient to support the continuous operation of the AI model currently.
[0046] In a specific example, the optimization of the AI model is as follows: Obtain each data in the AI model , record the floating-point range of each data as , and record the floating-point range of each data in the optimized AI model as , set the quantization step size as , according to the calculation formula: Analyze and obtain each optimized data , where Is the number corresponding to each data in the AI model, , is an arbitrary integer greater than 2, represents taking the integer of the data with the number .
[0047] It should be noted that by taking the integer of each data in the AI model, the storage space and computational complexity of the AI model are reduced.
[0048] In a specific example, the analysis obtains an adaptive algorithm, and the specific process is as follows: A1. Based on the hardware resource information of the user device, obtain the available CPU resource amount, available GPU resource amount, available memory resource amount, and network status of the user device. Based on the basic information of each inference path, obtain the path number, CPU resource occupancy, GPU resource occupancy, memory resource occupancy, and network dependency status of each inference path.
[0049] A2. Traverse each inference path. When the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy of a certain inference path are all less than or equal to the available CPU resource amount, available GPU resource amount, and available memory resource amount of the user device, and the network dependency status of this inference path conforms to the network status of the user device, mark this inference path as an available inference path. Accordingly, obtain all available inference paths in the user device, denoted as each available inference path; add the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy in each available inference path to obtain the total resource occupancy of each available path.
[0050] A3. After normalizing the total resource occupancy and inference accuracy of each available inference path, they are respectively denoted as and , where is the number corresponding to each available inference path, , is an arbitrary integer greater than 2. According to the calculation formula: analyze and obtain the priority selection index of each available inference path, where and respectively represent the weight factor corresponding to the total resource occupancy of the inference path and the weight factor corresponding to the inference accuracy.
[0051] It should be noted that the network status is divided into disconnected and connected, and the network dependency status is divided into dependent on the network and independent of the network.
[0052] It should be noted that the network dependency status of the inference path conforms to the network status of the user device, indicating that the network dependency status of the inference path is connected and the network status of the user device is connected, or the network dependency status of the inference path is disconnected and the network status of the user device is in any state.
[0053] It should be noted that the inference accuracy rates of all available inference paths are obtained from the cloud.
[0054] It should be noted that , , . Among them and The setting method of is the same as that of , so it will not be elaborated here.
[0055] In a specific example, the process of selecting the optimal inference path is as follows: Sort the priority selection indices of all available inference paths in descending order, construct a priority selection table of available inference paths for the AI model, and record the available inference path corresponding to the maximum priority selection index as the optimal inference path.
[0056] It should be noted that the optimal inference path is updated in real time according to the available CPU resources, available GPU resources, available memory resources, and network status of the user device.
[0057] Step 3: Task matching: Store each executed task of the user device in the cloud. When a task to be executed appears on the user device, match the task to be executed with each executed task in the cloud to obtain the matching task of the task to be executed, and then directly call the AI model and the optimal inference path of the matching task to process the task to be executed.
[0058] In a specific example, the process of matching the task to be executed with each executed task in the cloud to obtain the matching task of the task to be executed is as follows: Obtain the function description set and the resource requirement type set of the task to be executed from the user device; Obtain the function description set and the resource requirement type set of each executed task from the cloud, where is the number corresponding to each executed task, , is an arbitrary integer greater than 2. According to the calculation formula: Analyze and obtain the fitness of the task to be executed and each executed task, where and They respectively represent the weight factor corresponding to function matching when the client device executes a task and the weight factor corresponding to resource requirement type matching.
[0059] Compare the fitness of the to-be-executed task obtained with the fitness of each executed task with the set task fitness threshold. When the fitness of the to-be-executed task and a certain executed task is greater than or equal to the set task fitness threshold, it is determined that the to-be-executed task and the executed task match successfully; otherwise, it is determined that the to-be-executed task and the executed task do not match successfully. Based on this, obtain each executed task that matches the to-be-executed task successfully, and record the executed task with the highest fitness with the to-be-executed task as the matching task of the to-be-executed task.
[0060] It should be noted that , , . Among them and are set in the same way as , , so they will not be elaborated here.
[0061] An AI model dynamic optimization and adaptive inference method based on the client provided by this application analyzes the hardware resources and operating environment of the client device, then selects a suitable AI model to be deployed to the client device, and then monitors the operating state information of the client device in real time, so as to dynamically optimize and adjust the AI model to achieve the purpose of efficient inference. At the same time, according to the hardware resource information of the client device, the resource occupancy of each inference path, and the inference accuracy, an adaptive algorithm is analyzed. According to this adaptive algorithm, the client device can be made to adjust the optimal inference path in real time, so as to ensure the best inference performance and efficiency under limited resource conditions. Finally, store each processed task and its corresponding AI model and optimal inference path in the cloud for subsequent direct call.
[0062] The above content is only an example and illustration of the concept of this application. Those skilled in the art of this technology make various modifications or supplements to the described specific embodiments or use similar methods to replace them. As long as they do not deviate from the concept of the invention or exceed the scope defined by this application, they should all fall within the protection scope of this application.
Claims
1. A client-based AI model dynamic optimization and adaptive reasoning method, characterized in that: include: Step 1: Model selection: Analyze the hardware resource information and operating environment information of the user-end device, and then select the AI model to deploy to the user-end device; The hardware resource information of the user terminal device includes the available CPU resource, available GPU resource, available memory resource and network status of the user terminal device, and the operating environment information includes the battery power and network bandwidth of the user terminal device; The specific analysis process of analyzing the hardware resource information and operating environment information of the user terminal device is as follows: Based on the hardware resource information of the user-end device, the available CPU resources, GPU resources and memory resources of the user-end device are extracted, and then the hardware resource evaluation coefficient of the user-end device is obtained through analysis. ; Based on the operating environment information of the user-end device, the battery power and network bandwidth of the user-end device are extracted, and then the operating environment evaluation coefficient of the user-end device is obtained through analysis. ; Comprehensive user equipment hardware resource evaluation coefficient and operating environment assessment coefficient , according to the calculation formula: Analyze and obtain the available resource evaluation coefficient of the user terminal equipment ,in and They are respectively represented as the weight factor corresponding to the hardware resource evaluation coefficient of the user terminal device and the weight factor corresponding to the operating environment evaluation coefficient; The specific process of selecting an AI model to be deployed to a user-end device is as follows: Analyze the resource demand assessment coefficient of each AI model when it is deployed on the user-side device based on the basic operation information of the AI model, compare the resource demand assessment coefficient of each AI model when it is deployed on the user-side with the available resource assessment coefficient of the user-side device, and then obtain the ratio of the resource demand assessment coefficient of each AI model when it is deployed on the user-side to the available resource assessment coefficient of the user-side device, obtain each AI model corresponding to the ratio less than 1, record it as each selectable AI model, and obtain the inference accuracy of each selectable AI model from the cloud, arrange the inference accuracy of each selectable AI model in order from large to small, and select the AI model with the highest inference accuracy to deploy to the user-side device; Step 2: Model and reasoning path optimization: Real-time monitoring and analysis of the operating status information of the user-end device, and then optimization of the AI model, and analysis to obtain an adaptive algorithm to select the optimal reasoning path; The specific process of real-time monitoring and analyzing the operation status information of the user terminal equipment is as follows: The CPU temperature, CPU load and battery power consumption per unit time of the user-end device at the current moment are obtained from the monitored operation data of the user-end device, and recorded as , and , according to the calculation formula: Analyze and obtain the current operating status evaluation coefficient of the user terminal equipment ,in , and It is expressed as the standard value of CPU temperature, CPU load and battery power consumption per unit time of the user-side device at the current moment. , and They are respectively represented as the weight factor corresponding to the CPU temperature of the user terminal device at the current moment, the weight factor corresponding to the CPU load, and the weight factor corresponding to the battery power consumption per unit time; Compare the operation status evaluation coefficient of the user terminal device with the threshold of the operation device evaluation coefficient of the user terminal device, and when the operation status evaluation coefficient of the user terminal device is greater than or equal to the threshold of the operation status evaluation coefficient of the user terminal device, optimize the AI model, otherwise do not optimize the AI model; Step 3: Task matching: Store all executed tasks of the user-side device in the cloud. When a task to be executed appears on the user-side device, match the task to be executed with the executed tasks in the cloud, obtain the matching task of the task to be executed, and then directly call the AI model and optimal reasoning path of the matching task to process the task to be executed.
2. According to the client-based AI model dynamic optimization and adaptive reasoning method according to claim 1, it is characterized in that: The analysis obtains the hardware resource evaluation coefficient of the user terminal device and operating environment assessment coefficient , the specific analysis process is as follows: The available CPU resources, GPU resources, and memory resources of the client device are normalized and recorded as , and , according to the calculation formula: Analyze and obtain the hardware resource evaluation coefficient of the user-side device ,in , and They represent the weight factors corresponding to the available CPU resources, the available GPU resources, and the available memory resources respectively; similarly, the operating environment evaluation coefficient of the user-side device can be obtained by analyzing the battery power and network bandwidth of the user-side device. .
3. According to claim 2, a client-based AI model dynamic optimization and adaptive reasoning method is characterized in that: The specific process of optimizing the AI model is as follows: Get the data in the AI model , the floating point range of each data is recorded as , and record the floating point range of each data in the optimized AI model as , set the quantization step size to , according to the calculation formula: Analyze the optimized data ,in It is the number corresponding to each data in the AI model. , is any integer greater than 2, Indicates that the pair number is The data is rounded.
4. According to claim 3, a client-based AI model dynamic optimization and adaptive reasoning method is characterized in that: The analysis results in an adaptive algorithm, the specific process is as follows: A1. Based on the hardware resource information of the user-end device, obtain the available CPU resource amount, available GPU resource amount, available memory resource amount and network status of the user-end device, and based on the basic information of each reasoning path, obtain the path number, CPU resource occupancy, GPU resource occupancy, memory resource occupancy and network dependency status of each reasoning path; A2. Traverse each reasoning path. When the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy of a certain reasoning path are all less than or equal to the available CPU resource, available GPU resource, and memory resource occupancy of the user-end device, and the network dependency state of the reasoning path conforms to the network state of the user-end device, record the reasoning path as an available reasoning path, and thereby obtain all available reasoning paths in the user-end device, which are recorded as each available reasoning path; add the CPU resource occupancy, GPU resource occupancy, and memory resource occupancy in each available reasoning path to obtain the total resource occupancy of each available path; A3. The total resource usage and inference accuracy of each available inference path are normalized and recorded as and ,in is the number corresponding to each available reasoning path, , is any integer greater than 2, according to the calculation formula: Analyze and obtain the preference index of each available reasoning path ,in and They are respectively expressed as the weight factor corresponding to the total resource usage of the inference path and the weight factor corresponding to the inference accuracy.
5. According to claim 4, a client-based AI model dynamic optimization and adaptive reasoning method is characterized in that: The specific process of selecting the optimal reasoning path is as follows: Sort the priority index of each available reasoning path from large to small, build a priority table of available reasoning paths for the AI model, and record the available reasoning path corresponding to the maximum value of the priority index as the optimal reasoning path.
6. According to claim 5, a client-based AI model dynamic optimization and adaptive reasoning method is characterized in that: The specific process of matching the tasks to be executed with the executed tasks in the cloud to obtain the matching tasks of the tasks to be executed is as follows: Get the function description set of the tasks to be executed from the user end device and resource requirement type collection ; Get the functional description set of each executed task from the cloud and resource requirement type collection ,in is the number corresponding to each executed task, , is any integer greater than 2, according to the calculation formula: Analyze the compatibility between the tasks to be executed and the tasks that have been executed ,in and They respectively represent the weight factor corresponding to the function matching when the user terminal device performs the task and the weight factor corresponding to the resource requirement type matching; The obtained fitness of the task to be executed and each executed task is compared with the set task fitness threshold. When the fitness of the task to be executed and a certain executed task is greater than or equal to the set task fitness threshold, it is judged that the task to be executed and the executed task are matched successfully. Otherwise, it is judged that the task to be executed and the executed task are matched unsuccessfully. Based on this, each executed task that successfully matches the task to be executed is obtained, and the executed task with the highest fitness with the task to be executed is recorded as the matching task of the task to be executed.
Citation Information
Patent Citations
Multi-dimensional parallel processing method, system and device based on artificial intelligence, and readable storage medium
CN114035936A
Terminal and method for evaluating and testing ai task supporting capability of terminal
CN112204532A
Artificial intelligence AI model evaluation method, system and device
CN112508044A