Method, terminal device, server and system for deploying artificial intelligence model
The terminal device sends hardware and user preference information to the server and obtains matching operator databases, solving the problem of limited performance of AI models in the prior art, and achieving efficient AI model deployment and resource savings.
Patent Information
- Application Number
- CN202410028176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-01-08
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, the performance of the AI model of the terminal device is limited by different hardware performance, application scenarios and user preferences, resulting in the unification of the operator library that cannot match the characteristics of each terminal device, and the performance is poor.
The terminal device sends hardware information, resource information and user preference information to the server. The server obtains the matching operator library based on this information and returns it to the terminal device. The terminal device deploys the AI model based on the operator library.
It realizes the accuracy of issuing operator libraries to terminal devices, improves the performance of terminal devices' deployment of AI models, and saves the resources of terminal devices.
Smart Images

Figure CN120085875A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI), and in particular, to a method for deploying an AI model, a terminal device, a server, and a system. Background Art
[0002] AI is a science and technology for researching, developing, implementing, and applying intelligence, aiming to simulate human cognitive abilities through machines and enable machines to possess a certain degree of human intelligence. For example, AI can be used in various fields such as image recognition, natural language processing, speech processing, human-computer dialogue, and art creation.
[0003] In the prior art, a server can build an AI model based on the hardware performance of general and popular terminal devices and deploy the AI model to multiple terminal devices so that the AI model can run on each terminal device. However, due to the different hardware performances, application scenarios, and users of different terminal devices, the performance of the AI model in the terminal device is very limited. Summary of the Invention
[0004] In view of this, this application provides a method for deploying an artificial intelligence model, a terminal device, a server, and a system, which realizes the accuracy of sending an operator library to a terminal device, improves the performance of deploying an AI model on the terminal device, and saves the resources of the terminal device.
[0005] To achieve the above object, in a first aspect, an embodiment of this application provides a method for deploying an artificial intelligence model. The method includes: a terminal device sends hardware information, resource information, and first preference information of the terminal device to a server, where the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence (AI) model preferred by the user of the terminal device; the server obtains a first operator library matching the terminal device based on the hardware information, the resource information, and the first preference information, where the first operator library includes at least one operator; the server returns the first operator library to the terminal device; the terminal device deploys the AI model based on the first operator library.
[0006] In an embodiment of the present application, the server may obtain first preference information, hardware information, and resource information of a terminal device. The first preference information may indicate at least one AI model preferred by the user of the terminal device. The hardware information may indicate the hardware performance of the terminal device. The resource information may indicate the status of the device resources of the terminal device. Therefore, obtaining a first operator library from the server based on the first preference information, hardware information, and resource information can make the obtained first operator library match the hardware performance of the terminal device, the status of the device resources, and the user's preference for using AI models of the terminal device, achieving the accuracy of issuing the operator library to the terminal device and improving the performance of deploying the AI model on the terminal device. Since there is no need for the terminal device to deploy an optimization tool for searching operators and a large number of operator libraries are used by a small number of users, the device resources for storing the optimization tool and a large number of operator libraries on the terminal device are saved, and the resources for the terminal device to run the optimization tool to search for operators are also saved.
[0007] In some embodiments, the hardware information may include at least one of the type of the terminal device, the version of the terminal device, the type of at least one hardware device (such as a chip, a processor, a memory) in the terminal device, the version of at least one hardware device in the terminal device, the platform of at least one hardware device in the terminal device, the general-purpose level of at least one hardware device, and the like.
[0008] In some embodiments, the resource information may include at least one of the computing speed of the processor of the terminal device, the power consumption of the terminal device, the power consumption of at least one hardware device in the terminal device, the remaining storage space size of the memory, and the like.
[0009] In some embodiments, the method further includes: the terminal device sending user data of the terminal device to the server; the server obtaining second preference information based on the user data; the server returning the second preference information to the terminal device; the terminal device obtaining third preference information based on the user data; and the terminal device generating the first preference information based on the second preference information and the third preference information.
[0010] In some embodiments, the user data may include the historical record of the user using at least one AI model and / or the historical record of the user using at least one application program. Based on the historical record of the user using at least one AI model and / or the historical record of the user using at least one application program, information such as the frequency and duration of the user using at least one AI model can be determined, and further, the preference degree of the user for at least one AI model can be obtained, that is, the first preference information can be obtained.
[0011] In some embodiments, the method further includes: the terminal device obtaining third preference information based on the user data of the terminal device; the terminal device sending the user data and the third preference information to the server; the server obtaining second preference information based on the user data; and the server generating the first preference information based on the second preference information and the third preference information.
[0012] Both the terminal device and the server can learn the terminal device's preference for the AI model based on the user data of the terminal device, and obtain the second preference information and the third preference information respectively. Then, the second preference information and the third preference information are fused to obtain the first preference information, which improves the problem of long-tail users in the model when the server learns the user's preference for the AI model, and improves the overfitting problem caused by less user data on the terminal device, thereby improving the accuracy of learning the user's preference for the AI model.
[0013] In some embodiments, there may be at least one optimization tool in the server. Through the at least one optimization tool, a first operator library matching the terminal device is obtained based on the hardware information, the resource information, and the first preference information. In some embodiments, the optimization tool may include a TBE search tool, a cube computing engine (CCE) search tool, a CPU search tool, a heterogeneous search tool, a function flow template library (FFTL) search tool, or a quantization tool. The TBE search tool can be used to generate constant-quantized NPU operators. The CCE search tool can be used to search the output knowledge base and generate constant-quantized NPU operators based on CCE general operators. The CPU search tool can be used to search for constant-quantized CPU operators. The heterogeneous search tool can generate NPU operators and CPU operators based on a preset heterogeneous strategy. In some embodiments, the first operator library may include at least one of the TBE operators generated by the TBE search tool, the CCE operators and / or NPU operators generated by the CCE search tool, the CPU operators generated by the CPU search tool, the NPU operators and / or CPU operators generated by the heterogeneous search tool, the operators generated by the FFTL search tool, and the operators generated by the quantization tool.
[0014] In some embodiments, the server obtains second preference information based on the user data, including: the server inputs the user data into a first model to obtain the second preference information output by the first model; the terminal device obtains third preference information based on the user data, including: the terminal device inputs the user data into a second model to obtain the third preference information output by the second model, and the second model is obtained by compressing the first model.
[0015] In some embodiments, the structure of the second model is simpler than that of the first model, and / or the running efficiency of the second model is higher than that of the first model, and / or the second model includes fewer parameters than the first model. Therefore, the first model can be called a personalized large model, and the second model can be called a personalized small model.
[0016] In some embodiments, the method further includes: the server obtains the first model based on at least two pieces of user data;
[0017] The server compresses the first model to obtain the second model; the server returns the second model to the terminal device.
[0018] In some embodiments, the methods of compressing the first model include distillation, pruning, or quantization. Distillation means using a larger first model to guide the training of a smaller second model, and the effect of the trained second model can be close to that of the first model. Pruning means pruning the network structure with a lower degree of importance in the first model to obtain a second model with a simpler network structure. Quantization means reducing the data calculation precision in the first model to obtain a second model with a faster running speed.
[0019] In some embodiments, the first preference information includes a first vector, the first vector includes at least one element, each element in the at least one element corresponds to an AI model, and the value of each element indicates the user's preference degree for the AI model corresponding to the element.
[0020] In some embodiments, the second preference information includes a second vector, the third preference information includes a third vector, and the terminal device generates the first preference information based on the second preference information and the third preference information, including: the terminal device adds the product of the second vector and a first coefficient and the product of the third vector and a second coefficient to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.
[0021] In some embodiments, the terminal device deploys the AI model based on the first operator library, including: the terminal device compiles the AI model in the intermediate representation (IR) format based on the first operator library to obtain the AI model in the executable file format; and / or, the terminal device runs the AI model in the executable file format.
[0022] In some embodiments, the terminal device may compile an AI model in the executable file format based on at least one operator in the first operator library, so as to mount optimization parameters such as quantization weights, fused operators, and overall fusion parameters that match the user preferences, hardware performance of the device, and status of the device resources of the terminal device at the compilation state. In some embodiments, the terminal device may load and run the AI model in the executable file format based on at least one operator in the first operator library, so as to run the AI model based on the optimization parameters that match the user preferences, hardware performance of the device, and status of the device resources of the terminal device.
[0023] In some embodiments, the terminal device further includes a pre-set second operator library, and the second operator library includes at least one operator. The terminal device deploys the AI model based on the first operator library, including: the terminal device deploys the AI model based on the first operator library and the second operator library.
[0024] In a second aspect, an embodiment of the present application provides a method for deploying an artificial intelligence model, which is applied to a terminal device. The method includes: sending the hardware information, resource information, and first preference information of the terminal device to a server, where the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence (AI) model of the user preferences of the terminal device; obtaining a first operator library returned by the server, where the first operator library includes at least one operator; and the terminal device deploys the AI model based on the first operator library.
[0025] In some embodiments, the method further includes: sending the user data of the terminal device to the server; obtaining second preference information returned by the server based on the user data; obtaining third preference information based on the user data; and generating the first preference information based on the second preference information and the third preference information.
[0026] In some embodiments, the first preference information includes a first vector, the first vector includes at least one element, each element in the at least one element corresponds to an AI model, and the value of each element indicates the preference degree of the user for the AI model corresponding to the element.
[0027] In some embodiments, the second preference information includes a second vector, the third preference information includes a third vector, and generating the first preference information based on the second preference information and the third preference information includes: adding the product of the second vector and a first coefficient and the product of the third vector and a second coefficient to obtain a first vector, where the sum of the first coefficient and the second coefficient is 1.
[0028] In a third aspect, an embodiment of the present application provides a method for deploying an artificial intelligence model, which is applied to a server. The method includes: obtaining hardware information, resource information, and first preference information of a terminal device, where the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence (AI) model preferred by the user of the terminal device; obtaining a first operator library that matches the terminal device based on the hardware information, the resource information, and the first preference information, where the first operator library includes at least one operator; and returning the first operator library to the terminal device, where the first operator library is used for the terminal device to deploy the AI model.
[0029] In some embodiments, the method further includes: obtaining user data and third preference information of the terminal device; obtaining second preference information based on the user data; and generating the first preference information based on the second preference information and the third preference information.
[0030] In some embodiments, the method further includes: obtaining a first model based on at least two pieces of user data, where the first model is used for the server to obtain the third preference information; performing compression processing on the first model to obtain a second model, where the second model is used for the terminal device to obtain the second preference information; and returning the second model to the terminal device.
[0031] In a fourth aspect, an embodiment of the present application provides a system, which includes a server and a terminal device; the terminal device is configured to send the hardware information, resource information, and first preference information of the terminal device to the server, where the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence (AI) model preferred by the user of the terminal device; and deploy the AI model based on a first operator library, where the first operator library includes at least one operator; the server is configured to obtain a first operator library that matches the terminal device based on the hardware information, the resource information, and the first preference information; and return the first operator library to the terminal device.
[0032] Fifth aspect, an embodiment of the present application provides a device for deploying an artificial intelligence model. The device has the function of implementing the behaviors of the terminal device in the above aspects and the possible implementation manners of the above aspects. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a transceiver module or unit, a processing module or unit, an acquisition module or unit, etc.
[0033] Sixth aspect, an embodiment of the present application provides a device for deploying an artificial intelligence model. The device has the function of implementing the behaviors of the server in the above aspects and the possible implementation manners of the above aspects. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a transceiver module or unit, a processing module or unit, an acquisition module or unit, etc.
[0034] Seventh aspect, an embodiment of the present application provides a terminal device, including: a memory and a processor. The memory is used for storing a computer program; the processor is used for executing the method described in any one of the above second aspects when calling the computer program.
[0035] Eighth aspect, an embodiment of the present application provides a chip system. The chip system includes a processor, and the processor is coupled to a memory. The processor executes the computer program stored in the memory to implement the method described in any one of the above second aspects.
[0036] Wherein, the chip system can be a single chip or a chip module composed of multiple chips.
[0037] Ninth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above second aspects is implemented.
[0038] Tenth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the method described in any one of the above second aspects.
[0039] Eleventh aspect, an embodiment of the present application provides a server, including: a memory and a processor. The memory is used for storing a computer program; the processor is used for executing the method described in any one of the above third aspects when calling the computer program.
[0040] Twelfth aspect, an embodiment of the present application provides a chip system. The chip system includes a processor, and the processor is coupled to a memory. The processor executes the computer program stored in the memory to implement the method described in any one of the above third aspects.
[0041] Among them, the chip system may be a single chip or a chip module composed of multiple chips.
[0042] In a thirteenth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above third aspects is implemented.
[0043] In a fourteenth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a server, the server is caused to execute the method described in any one of the above third aspects.
[0044] It can be understood that the beneficial effects of the above second aspect to fourteenth aspect can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0045] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0046] Figure 2 It is a structural block diagram of a system provided by an embodiment of the present application;
[0047] Figure 3 It is a schematic flow diagram of a method for deploying an AI model provided by an embodiment of the present application;
[0048] Figure 4 It is a schematic diagram of an AI model deployment process provided by an embodiment of the present application;
[0049] Figure 5 It is a schematic flow diagram of a method for learning user preference information provided by an embodiment of the present application. Detailed Embodiments
[0050] The method for deploying a machine learning model provided by an embodiment of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The specific type of the electronic device is not limited in the embodiments of the present application.
[0051] Figure 1It is a schematic structural diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 may include a processor 110, a memory 120, a communication module 130, etc.
[0052] Among them, the processor 110 may include one or more processing units, and the memory 120 is used to store program codes and data. In the embodiment of the present application, the processor 110 can execute the computer execution instructions stored in the memory 120 to control and manage the actions of the electronic device 100. For example, the processor 110 may include a system-on-a-chip (SOC). For example, the processor 110 may include a graphics processing unit (GPU), a neural-network processing unit (NPU), or a central processing unit (CPU). In some embodiments, the processor 110 may include an AI core and an AI processor. The AI core can be used to run operators of specific algorithms, such as matrix, vector, and scalar calculation-intensive operators. The AI processor can be used as a supplement to the AI processor to run other operators.
[0053] The communication module 130 can be used for communication between various internal modules of the electronic device 100, or communication between the electronic device 100 and other external electronic devices, etc. Exemplarily, if the electronic device 100 communicates with other electronic devices in a wired connection manner, the communication module 130 may include an interface, etc., such as a USB interface. The USB interface can be an interface that conforms to the USB standard specification, specifically a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other electronic devices.
[0054] Alternatively, the communication module 130 may include audio devices, radio frequency circuits, Bluetooth chips, wireless fidelity (Wi-Fi) chips, near-field communication (NFC) modules, etc., and can implement the interaction between the electronic device 100 and other electronic devices in a variety of different ways.
[0055] The wireless communication function of the electronic device 100 can be implemented by at least one antenna, a mobile communication module, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.
[0056] Among them, the antenna is used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0057] Optionally, the electronic device 100 may further include a display screen 140, and the display screen 140 can display images, videos, etc. in the human-computer interaction interface. The electronic device 100 can implement the display function through a microprocessor for image processing, a display screen, an application processor, etc.
[0058] The display screen 140 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, and N is a positive integer greater than 1.
[0059] The electronic device 100 can implement the shooting function through a camera, a video codec, a GPU, the display screen 140, an application processor, etc.
[0060] Optionally, the electronic device 100 may further include peripheral devices 150, such as a mouse, a keyboard, a speaker, a microphone, etc.
[0061] Optionally, the electronic device 100 may further include a charging management module, a power management module, a battery, a key, an indicator, and one or more SIM card interfaces, etc., and the embodiments of the present application do not make any restrictions on this.
[0062] It should be understood that in addition to Figure 1Except for the various components or modules listed, the embodiments of the present application do not specifically limit the structure of the electronic device 100. In some other embodiments of the present application, the electronic device 100 may further include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0063] First, some of the terms involved in the embodiments of the present application are explained:
[0064] AI is a science and technology that studies, develops, implements, and applies intelligence, aiming to simulate human cognitive abilities through machines and enable machines to possess a certain degree of human intelligence. For example, AI can be used to achieve various capabilities in fields such as image recognition, natural language processing, speech processing, human-machine dialogue, and artistic creation.
[0065] An AI model is a mathematical model that analyzes, processes, predicts, and optimizes data with certain regularity and predictability. An AI model corresponding to the requirements of a specific business or capability field can be constructed, and this AI model has the ability to process this specific business. Exemplarily, according to the business type of the AI model, an optical character recognition (OCR) model can recognize the text in an image, a natural language understanding (NLU) model can understand the semantics of human language and recognize the user's intention, a natural language generation (NLG) model can generate text information that humans can understand, a text-to-speech (TTS) model can generate voice information based on the text information, a face recognition model can recognize the face included in the image, and a super-resolution model can improve the resolution of the input image to achieve high definition.
[0066] Exemplarily, the AI model may include a linear model, a decision tree model, an ensemble model, a neural network model, and a support vector machine model. The linear model is a simple model that can divide or predict data with a straight line or a hyperplane. Common linear models include linear regression, logistic regression, etc. The decision tree model is a tree-structured model that divides the dataset into smaller subsets until each subset contains data points of a single class, such as classification tree models and regression tree models, etc. The ensemble model is a model that improves the prediction performance by combining multiple models, such as random forest models, gradient boosting tree models, etc. The neural network model is a model based on the biological nervous system that establishes complex mapping relationships through the connections between multiple neurons, such as multi-layer perceptron models, convolutional neural network models, recurrent neural network models, etc. The support vector machine model is a model based on maximum margin classification that maps data points into a high-dimensional space and finds the maximum margin hyperplane to divide the data, such as linear support vector machine models, non-linear support vector machine models, etc.
[0067] In some embodiments, the AI model may include a computational graph (graph) and weights.
[0068] The computational graph is used to represent the operation process of the algorithm and is a method of formalizing the algorithm. The computational graph includes multiple nodes, which are connected by directed edges, and each node represents an operator. The input edges entering the node represent the input data of the operator corresponding to the node, and the output edges leaving the node represent the output data of the operator corresponding to the node. The operation process represented by the computational graph can be a model inference process or a model training process. Save each operator according to the connection method of the directed edges in the computational graph, as well as the weights involved in each operator, so as to realize the saving of the AI model.
[0069] The operator is the basic component that makes up the AI model and can implement one or more operation operations in the AI model. For example, the convolution operator can be used for feature extraction and mapping of data; the activation function operator can be used to introduce non-linear factors to improve the expression ability of the AI model and solve problems that cannot be solved by linear models; the loss function operator can be used to measure the error between the prediction result and the true result of the AI model, and then train, evaluate, or optimize the AI model according to this error. In some embodiments, the operators in the AI model may correspond to layers in the AI model. In some embodiments, the parameters of the operator may include the input data of the operator, the output data of the operator, and the elements in the operator (such as the size of the convolution kernel in the convolution operator).
[0070] The weights are used to represent the data that the operator needs to use during execution.
[0071] A task unit is a smaller scheduling unit obtained by decomposing an operator. It can achieve scheduling between and within operators, break the operator boundary, and allow fine-grained scheduling of computations onto hardware devices. In some embodiments, an operator can be divided into multiple task units.
[0072] An operator kernel is the specific implementation of an operator based on the hardware device that executes the operator. The operator kernel can be obtained by compiling the operator. The operator kernel implements the computational logic of specific task units and determines the total number of task units. In some embodiments, when the same operator is compiled for different hardware devices (such as a CPU and an NPU), the obtained operator kernels are also different. In some embodiments, an operator can be implemented as one or more operator kernels. In some embodiments, the operator kernel can include files in the *.so format and files in the *.json format.
[0073] The model format of an AI model can include the following several formats.
[0074] A framework model is an AI model in the model format output by a deep learning framework. Exemplarily, the deep learning framework can include a Hopfield neural network (HNN) / MindSpore. The framework model is independent of the hardware on which the AI model actually runs.
[0075] An intermediate representation (IR) model can represent the structure and computational process of an AI model in the form of a computational graph. The IR model is independent of the hardware on which the AI model actually runs.
[0076] An AI model in executable file format. The executable file is a file that can run on hardware. The executable file format is related to the hardware on which the AI model actually runs. Different types of hardware can respectively correspond to different executable files. For example, if the hardware on which the terminal device runs the AI model is a CPU, the executable file can correspond to that CPU.
[0077] In some embodiments, since the framework model and the IR model are independent of the hardware on which the AI model actually runs, and the executable file format is related to the hardware on which the AI model actually runs, the framework model and the IR model can also be high-level model formats, and the executable file format can be a low-level model format. After the AI model in the high-level model format is compiled by a compilation tool for the hardware on which the AI model runs, an AI model in the low-level model format can be obtained.
[0078] In some embodiments, the compilation tool may include a tensor virtual machine (TVM) or a tensor boost engine (TBE). It can be understood that in practical applications, other compilation tools may also be used to compile the AI model.
[0079] Taking TVM as an example, the compilation process of the compilation tool for the AI model may include: obtaining the AI model in the high-level model format and generating the computation graph of the AI model; optimizing the computation graph (for example, high-level data flow rewriting and / or operator-level optimization) to obtain the optimized computation graph; and compiling and generating an executable file that can be run on hardware devices such as the processor in the terminal device based on the optimized operator graph.
[0080] To facilitate the understanding of the technical solutions in the embodiments of the present application, the application scenarios of the embodiments of the present application will be introduced below.
[0081] To achieve the intelligence of the terminal device, various AI models can be deployed in the terminal device, and each AI model can have the ability to process at least one service. When deploying the AI model, how to improve the performance experience of the services processed by the AI is a very crucial issue. For example, for a terminal device with relatively poor hardware performance, by deploying and optimizing the AI model deployed on the terminal device, the performance of the AI model in processing related services on the terminal device can be improved, thereby significantly reducing the gap between the performance experience of the service on the terminal device and the performance experience of the service on other terminal devices with better hardware performance.
[0082] As can be seen from the foregoing, the operator is the basic component that constitutes the AI model. Therefore, according to the hardware information of the terminal device, the operator library of the AI model can be searched through an optimization tool. The operator library includes at least one operator, and the operator library can indicate the model structure and operation parameters of the AI model. By deploying the AI model according to the searched operator library, the deployed AI model can have a better model structure and model operation process.
[0083] In some embodiments, the optimization tool may include at least one of a TBE search tool, a CCE search tool, a CPU search tool, a heterogeneous search tool, an FFTL search tool, and a quantization tool. The TBE search tool can be used to generate constant-quantized NPU operators. The CCE search tool can be used to search the output knowledge base and generate constant-quantized NPU operators based on CCE general operators. The CPU search tool can be used to search for constant-quantized CPU operators. The heterogeneous search tool can generate NPU operators and CPU operators based on a preset heterogeneous strategy.
[0084] In some embodiments, an optimization tool is included in the server. The server can search for a general-purpose operator library according to the hardware performance of general and popular terminal devices through the optimization tool, and send the operator library to each terminal device, so that each terminal device can deploy an AI model based on the operator library. However, for different terminal devices, their product forms, product versions, application scenarios, application preferences, etc. are all different. Therefore, this method of uniformly sending a general-purpose operator library to terminal devices cannot generate an operator library suitable for each terminal device, and the performance of the deployed AI model is very poor.
[0085] In some embodiments, an optimization tool is included in the terminal device. The terminal device can search for an operator library matching the terminal device according to the hardware information of the terminal device. However, since the optimization tool needs to occupy the memory of the terminal device, and searching for the optimal operator through the optimization tool also requires memory and power consumption, for terminal devices with limited resources such as read-only memory (ROM) and random access memory (RAM), it is impossible to store and apply the optimization tool.
[0086] To solve at least some of the above technical problems, an embodiment of the present application provides another method for deploying an AI model. The server can search for an operator library matching the terminal device through an optimization tool according to the hardware information, user preference information, and resource information of the terminal device. The terminal device can dynamically manage the required operator library according to the user's usage habits and frequencies, minimizing the required space of the terminal device. When the terminal device is inferring, it is deployed online through the operator library matching the terminal device, achieving performance optimization during compilation and operation.
[0087] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0088] Please refer to Figure 2 , which is a schematic structural diagram of a system 200 provided by an embodiment of the present application. The system 200 may include a server 210 and a terminal device 220, and the server 210 and the terminal device 220 may be connected through a network. The server 210 and the terminal device 220 may be implemented by the electronic device 100 described above.
[0089] The server 210 may include an optimization tool. Through this optimization tool, the server 210 can search for a dedicated operator library that matches the terminal device 220 according to the relevant information of the terminal device 220, such as hardware information, user preference information, and resource information. The dedicated operator library may include at least one operator that matches the terminal device 220. The server 210 may publish the dedicated operator library so that the terminal device 220 can obtain the dedicated operator library.
[0090] The terminal device 220 includes an application, an AI operator library, and a processor.
[0091] The application may include or invoke an AI model. The terminal device 220 may include multiple applications. The same application may include or invoke multiple AI models with different business types, and different applications may also include or invoke the same business type of AI model. For example, the video application A may include a super-resolution model, the video application B may also include a super-resolution model, and the voice assistant may include an automatic speech recognition (ASR) model, an OCR model, etc.
[0092] The AI operator library may include a dedicated operator library that matches the terminal device 220. In some embodiments, the terminal device 220 may dynamically download the dedicated operator library from the server 210. In some embodiments, the terminal device 220 may also include a general operator library. The general operator library may include at least one operator, and for different terminal devices 220, the included general operator libraries may be the same.
[0093] The processor may be used to run one or more operators, thereby implementing the services processed by the application and the AI model.
[0094] Based on Figure 2 The system 200 shown. The server 210 may search for and publish a dedicated operator library that matches the terminal device 220 based on the relevant information such as the hardware information, user preference information, and resource information of the terminal device 220. The terminal device 220 may download the dedicated operator library from the server 210 and deploy the AI model based on the operators in the dedicated operator library, so as to achieve performance optimization during compilation and runtime.
[0095] Please refer to Figure 3 , which is a flowchart of a method for deploying an AI model provided by an embodiment of the present application. The method may be applied to Figure 2 The system shown. It should be noted that the method is not based on Figure 3With the following specific order as a limitation, it should be understood that in other embodiments, the order of some steps of the method can be mutually exchanged according to actual needs, or some of the steps can also be omitted or deleted. The method includes the following steps:
[0096] S301, the terminal device and / or the server perform personalized learning on the user preferences of the terminal device to obtain first preference information.
[0097] Among them, the first preference information can be used to indicate at least one AI model of the user preferences of the terminal device. In some embodiments, the first preference information can also be used to indicate the preference degree of the user of the terminal device for each AI model in at least one AI model.
[0098] In some embodiments, the terminal device and / or the server can determine the first preference information based on the user data of the terminal device. The user data can be obtained by the terminal device during the user's use of the terminal device. In some embodiments, the user data can include the historical records of the user using at least one AI model and / or the historical records of the user using at least one application. Based on the historical records of the user using at least one AI model and / or the historical records of the user using at least one application, information such as the frequency and duration of the user using at least one AI model can be determined, and further, the preference degree of the user for at least one AI model can be obtained, that is, the first preference information can be obtained. Among them, the preference degree of the user for the AI model can be positively correlated with the frequency and duration of the user using the AI model.
[0099] In some embodiments, the terminal device sends the user data of the terminal device to the server, the server obtains second preference information based on the user data, the server returns the second preference information to the terminal device, the terminal device obtains third preference information based on the user data, and the terminal device generates the first preference information based on the second preference information and the third preference information. Among them, the second preference information can indicate at least one AI model preferred by the user learned by the server, and the third preference information can be at least one AI model preferred by the user learned by the terminal device. That is, the first preference information can be jointly determined by the second preference information and the third preference information obtained by the terminal device and the server respectively based on the user data of the terminal device, improving the accuracy of determining the first preference information.
[0100] In some embodiments, the server includes a first model, and the server can input the user data of the terminal device into the first model to obtain the second preference information output by the first model.
[0101] Among them, the first model can receive the input user data and output corresponding preference information. In some embodiments, the server can obtain the user data of at least two terminal devices, and train and obtain the first model based on the obtained at least two user data. Alternatively, in some other embodiments, the server can obtain a trained first model. In some embodiments, if the server obtains new user data, it can also perform incremental training on the first model based on the new user data, thereby improving the accuracy of the first model.
[0102] In some embodiments, the terminal device includes a second model, and the terminal device can input the user data of the terminal device into the second model to obtain the third preference information output by the second model.
[0103] Among them, the second model can receive the input user data and output corresponding preference information. In some embodiments, the terminal device can train and obtain the second model based on the user data of the terminal device. Alternatively, in some other embodiments, the terminal device can obtain a trained second model. Exemplarily, the terminal device can obtain the second model from the server. In some embodiments, if the terminal device obtains new user data, it can also perform incremental training on the second model based on the new user data, thereby improving the accuracy of the second model.
[0104] It can be understood that the first model and the second model can be the same or similar, and the manner in which the server obtains the first model can also be the same or similar to the manner in which the terminal device obtains the second model. In some embodiments, the structure of the second model is simpler than that of the first model, and / or the operating efficiency of the second model is higher than that of the first model, and / or the second model includes fewer parameters than the first model. Therefore, the first model can be called a personalized large model, and the second model can be called a personalized small model.
[0105] In some embodiments, the terminal device obtains the first preference information based on the user data of the terminal device. That is, the first preference information is determined only through the terminal device.
[0106] In some embodiments, the terminal device can receive the first preference information submitted by the user. Exemplarily, the terminal device can display an interface for receiving the first preference information to the user and receive the first preference information submitted by the user through this interface. Alternatively, the terminal device can receive a file recording the first preference information submitted by the user and obtain the first preference information from this file.
[0107] In some embodiments, the terminal device obtains third preference information based on the user data of the terminal device, the terminal device sends the user data and the third preference information to the server, the server obtains second preference information based on the user data, and the server generates first preference information based on the second preference information and the third preference information.
[0108] In some embodiments, the terminal device sends the user data of the terminal device to the server, and the server obtains first preference information based on the user data. That is, the first preference information is determined only by the server.
[0109] In addition, the manner in which the terminal device and the server perform personalized learning on the user preferences of the terminal device to obtain the first preference information can also refer to the detailed description in the following Figure 5 below.
[0110] S302, the terminal device obtains a first operator library from the server based on the first preference information, the hardware information, and the resource information.
[0111] The first preference information may indicate at least one AI model preferred by the user of the terminal device. The hardware information may indicate the hardware performance of the terminal device. The resource information may indicate the status of the device resources of the terminal device. Therefore, obtaining the first operator library from the server based on the first preference information, the hardware information, and the resource information can make the obtained first operator library match the hardware performance of the terminal device, the status of the device resources, and the user's preference for using the AI model of the terminal device, realizing the accuracy of issuing the operator library to the terminal device. Since there is no need for the terminal device to deploy an optimization tool for searching operators and a large number of operator libraries are used less by users, the device resources for storing the optimization tool and a large number of operator libraries in the terminal device are saved, and the resources for the terminal device to run the optimization tool to search for operators are also saved.
[0112] In some embodiments, the hardware information may include at least one of the type of the terminal device, the version of the terminal device, the type of at least one hardware device (such as a chip, a processor, a memory) in the terminal device, the version of at least one hardware device in the terminal device, the platform of at least one hardware device in the terminal device, the general-purpose level of at least one hardware device, and the like.
[0113] In some embodiments, the resource information may include at least one of the computing speed of the processor of the terminal device, the power consumption of the terminal device, the power consumption of at least one hardware device in the terminal device, the remaining storage space size of the memory, and the like.
[0114] In some embodiments, the terminal device may send first preference information, hardware information, and resource information to the server. Based on the first preference information, hardware information, and resource information, the server searches for a first operator library that matches the terminal device and sends the first operator library to the terminal device. It can be understood that if the first preference information is determined by the server, the terminal device may send the hardware information and resource information to the server.
[0115] In some embodiments, the server may search for a first operator library that matches the first preference information, hardware information, and resource information from multiple stored operator libraries based on at least one optimization tool. In some embodiments, the first operator library may include at least one operator among TBE operators generated by the TBE search tool, CCE operators and / or NPU operators generated by the CCE search tool, CPU operators generated by the CPU search tool, NPU operators and / or CPU operators generated by the heterogeneous search tool, operators generated by the FFTL search tool, and operators generated by the quantization tool. Among them, the server may obtain and store multiple operator libraries in advance. Exemplarily, the multiple operator libraries may include a general-purpose operator library, a CV-class operator library, an L0-level device operator library, and an operator library for users who only use video models. And it can be understood that the specific content of the multiple operator libraries included in the server in the embodiments of the present application is not limited.
[0116] S303, the terminal device deploys the AI model based on the first operator library.
[0117] When the terminal device obtains a first operator library that matches the first preference information, hardware information, and resource information of the terminal device from the server, it may deploy the AI model based on the first operator library, thereby realizing the optimized deployment of the AI model based on the user preferences of the terminal device, the hardware performance of the device, and the status of the device resources.
[0118] The AI model may be an AI model to be deployed. In some embodiments, when the terminal device obtains an application request to deploy an AI model, it deploys the AI model requested by the application based on the first operator library.
[0119] In some embodiments, the terminal device compiles an IR model (i.e., an AI model in IR format) based on a first operator library to obtain an AI model in executable file format; and / or, the terminal device can run the AI model in executable file format. In some embodiments, the terminal device can compile an AI model in executable file format based on at least one operator in the first operator library, so as to mount optimization parameters such as quantization weights, fused operators, and overall fusion parameters that match the user preferences, hardware performance of the device, and status of device resources at the compilation state. In some embodiments, the terminal device can load and run the AI model in executable file format based on at least one operator in the first operator library, so as to run the AI model based on optimization parameters that match the user preferences, hardware performance of the device, and status of device resources of the terminal device.
[0120] In some embodiments, the terminal device can compile the operators in the AI model into one or more operator cores and one or more task units based on the hardware device for executing operators in the terminal device, and generate code and instructions based on the one or more operator cores and the one or more task units, so as to obtain an AI model in executable file format.
[0121] Exemplarily, the manner in which the terminal device deploys the AI model based on the first operator library can be as Figure 4 shown. Among them, the first terminal 230 can be the terminal device held by the user.
[0122] The server 210 includes an optimization tool, and searches through the optimization tool to obtain a first operator library that matches the user preferences, hardware performance of the device, and status of device resources of the first terminal 230. The server 210 publishes the first operator library. In some embodiments, operator search and compilation, model conversion, recording weight rearrangement rules, and recording quantization parameters can be performed through the optimization tool to obtain the first operator library.
[0123] The first terminal 230 obtains and stores the first operator library from the server 210. The first terminal 230 can obtain a framework model through a deep learning framework, convert the framework model into an IR model, and compile the IR model based on the first operator library, so as to mount the rearrangement rules of the task units, operator cores, and weights carried by the first operator library to obtain an AI model in executable file format. The first terminal 230 can store the AI model in executable file format, and load and run the AI model in executable file format.
[0124] In some embodiments, before releasing the first operator library, the server 210 may also send the first operator library to the second terminal 240, so that relevant technical personnel can test the first operator library through the second terminal 240. Deploying the AI model based on the first operator library can determine the performance improvement of the first operator library on the deployed AI model. When it is determined that the first operator library can improve the performance of the deployed AI model, the server 210 may release the first operator library. Among them, the second terminal 240 may be a test machine held by relevant technical personnel, and the second terminal 240 may be similar to or the same as the first terminal device 230.
[0125] In some embodiments, the partial file structure of an AI model in an executable file format may be as follows:
[0126] / data / hiai / hash_id / # The hash_id is the hash value of the computation graph output by the IR model
[0127] |----main /
[0128] |----model_info.json# Configuration file of the AI model
[0129] |----xyz.omc# AI model without weights
[0130] |----subgraph_0 / # TVM operator
[0131] |----xxx.so# Library file
[0132] |----yyy.json# Weight conversion record file
[0133] |----subgraph_1 / # TBE operator
[0134] |----zzz.json# Weight conversion record file
[0135] |----subgraph_2 / # Processor Computing Language (CPUCL) operator
[0136] |----aaa.json# Weight conversion record file
[0137] |----bbb.so# Library file, including the core of the CPUCL operator library used
[0138] |----subgraph_3 / # CCE operator
[0139] |----ccc.json# Weight conversion record file
[0140] Among them, subgraph is the weight conversion configuration file of the sub-computation graph. Each AI model may include multiple subgraphs, and each subgraph corresponds to the weight conversion configuration file of the sub-computation graph of an operator of a certain type.
[0141] In some embodiments, the configuration file of the AI model may include at least one of the information such as the file name of the AI model, the version number, and the segmentation information of the sub-computation graph.
[0142] In some embodiments, the terminal device further includes a pre-set second operator library, and the second operator library may include at least one operator. The terminal device may deploy the AI model based on the first operator library and the second operator library. In some embodiments, the operators in the second operator library may be general-purpose operators.
[0143] In some embodiments, the differences between the first operator library and the second operator library may include:
[0144] Difference one: The first operator library includes dedicated operators (also referred to as specific operators) deployed based on the first preference information of the terminal device, the hardware information of the device, and the resource information. The second operator library may include general-purpose operators. In some embodiments, the first operator library may be the Figure 2 custom operator library mentioned above, and the second operator library may be the Figure 2 general operator library mentioned above.
[0145] Difference two: The first operator library is an operator library dynamically deployed based on the dynamic changes of the first preference information of the terminal device, the hardware information of the device, and the resource information. When at least one of the first preference information, the hardware information of the device, and the resource information changes, the operators included in the first operator library will also be updated. The second operator library will not be updated based on the changes of the first preference information, the hardware information of the device, and the resource information.
[0146] In an embodiment of the present application, the server may obtain first preference information, hardware information, and resource information of a terminal device. The first preference information may indicate at least one AI model preferred by the user of the terminal device, the hardware information may indicate the hardware performance of the terminal device, and the resource information may indicate the status of the device resources of the terminal device. Therefore, obtaining a first operator library from the server based on the first preference information, hardware information, and resource information can make the obtained first operator library match the hardware performance of the terminal device, the status of the device resources, and the user's preference for using the AI model of the terminal device, achieving the accuracy of issuing the operator library to the terminal device and improving the performance of deploying the AI model on the terminal device. Since there is no need for the terminal device to deploy an optimization tool for searching operators and a large number of operator libraries are used by a small number of users, the device resources for storing the optimization tool and a large number of operator libraries on the terminal device are saved, and the resources for the terminal device to run the optimization tool to search for operators are also saved.
[0147] Please refer to Figure 5 , which is a flowchart of a method for obtaining preference information provided by an embodiment of the present application. Among them, this method may be a detailed description of the foregoing S301. It should be noted that this method is not limited by Figure 5 the specific order described below. It should be understood that in other embodiments, the order of some steps of this method may be interchanged according to actual needs, or some of the steps may also be omitted or deleted. The method includes the following steps:
[0148] S501, the server trains a first model based on the user data of at least two terminal devices.
[0149] The user data of at least two terminal devices may include the historical records of the users of at least two terminal devices using AI models and / or the records of at least one application of the users of at least two terminal devices. The server trains and obtains the first model based on the user data of at least two terminal devices, that is, learns the preferences of group users for AI models.
[0150] In some embodiments, the server may separately receive the user data sent by at least two terminal devices. In some embodiments, the user data may be non-private data.
[0151] In some embodiments, the server may train the first model every first time period or every time the first number of user data is obtained to ensure the accuracy of the first model. It should be noted that the embodiments of the present application do not limit the ways of setting the first time period and the first number, nor the specific sizes of the first time period and the first number. Exemplarily, the first time period and the first number may be set or updated by those skilled in the relevant art in advance.
[0152] In some embodiments, the server may obtain the first model in other ways. Exemplarily, the server may obtain the trained first model, so S501 may be omitted.
[0153] S502. The server compresses the first model to obtain a second model.
[0154] Among them, compressing the first model, that is, miniaturizing the first model, can make the obtained second model run faster and consume fewer device resources. Since the device resources in the terminal device are less than those in the server, the server can compress the first model to obtain a second model that is convenient for deployment in the terminal device, so that the terminal device can also learn the user's preference for the AI model from the user data.
[0155] In some embodiments, the ways to compress the first model include distillation, pruning, or quantization.
[0156] Distillation, that is, using a larger first model to guide the training of a smaller second model, and the effect of the trained second model can be close to that of the first model.
[0157] Pruning, that is, pruning the network structure with relatively low importance in the first model to obtain a second model with a relatively simple network structure.
[0158] Quantization, that is, reducing the data calculation precision in the first model to obtain a second model with a faster running speed.
[0159] It can be understood that in practical applications, the first model can also be compressed in other ways.
[0160] In some embodiments, if the server updates the first model, it can update the second model deployed on the terminal device based on the updated first model.
[0161] In some embodiments, the terminal device may obtain the second model in other ways. Exemplarily, the terminal device may train and obtain the second model based on the user data of the terminal device, so S502 may be omitted.
[0162] S503. The server performs personalized learning based on the user preferences of the terminal device to obtain second preference information, and the terminal device performs personalized learning on the user preferences of the terminal device to obtain third preference information.
[0163] The terminal device can send the user data of the terminal device to the server. The server learns from the user data based on the first model to obtain second preference information, and the second preference information is the coarse-grained features learned by the first model. The terminal device can also learn from the user data based on the second model to obtain third preference information, and the third preference information is the fine-grained features learned by the second model. In some embodiments, the server can return the second preference information to the terminal device.
[0164] In some embodiments, the server can input the user data of the terminal device into the first model to obtain the second preference information output by the first model. In some embodiments, the second preference information can indicate the AI model preferred by the user based on the service type corresponding to the AI model and the application program to which the AI model belongs. That is to say, each AI model corresponds to a service type and an application program. Exemplarily, the super-resolution model belonging to video application A and the super-resolution model belonging to video application B can be used as two AI models, so as to respectively indicate the degree of preference of the user for the super-resolution model belonging to video application A and the degree of preference of the user for the super-resolution model belonging to video application B.
[0165] In some embodiments, the second preference information can include a second vector. The second vector includes at least one element. Each element in the at least one element corresponds to an AI model, and the value of each element indicates the degree of preference of the user for the AI model corresponding to the element.
[0166] For example, the second vector can be [0.1, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.2]. Among them, each element can correspond to an AI model, and the value of each element is greater than or equal to 0 and less than or equal to 1. The closer the value of each element is to 1, the more the user prefers to use the AI model corresponding to the element.
[0167] Alternatively, in some other embodiments, the second preference information can indicate the AI model preferred by the user based on the service type corresponding to the AI model. Exemplarily, the super-resolution model belonging to video application A and the super-resolution model belonging to video application B can be used as one AI model, so as to indicate the degree of preference of the user for the super-resolution model.
[0168] In some embodiments, the server may input the user data of the terminal device into a first model to obtain fourth preference information output by the first model, and the server compresses the fourth preference information to obtain second preference information. The fourth preference information may indicate the AI model preferred by the user based on the service type corresponding to the AI model, the application to which the AI model belongs, and more dimensions for differentiating the AI model, so as to more accurately indicate the user's preference for the AI model based on multiple dimensions. In some embodiments, the fourth preference information may indicate the AI model preferred by the user based on the device type where the AI model is located, the service type corresponding to the AI model, and the application to which the AI model belongs. In other words, each AI model corresponds to a terminal device, a service type, and an application. Exemplarily, the ASR model deployed in the in-vehicle unit and the ASR model deployed in the mobile phone may be used as two AI models to respectively indicate the degree of preference of the user for the ASR model deployed in the in-vehicle unit and the degree of preference of the user for the ASR model deployed in the mobile phone.
[0169] In some embodiments, the fourth preference information may include a fourth vector, and the fourth vector includes at least one element. Each element in the at least one element corresponds to an AI model, and the value of each element indicates the degree of preference of the user for the AI model corresponding to the element.
[0170] In some embodiments, the terminal device may input the user data of the terminal device into a second model to obtain third preference information output by the second model. In some embodiments, the third preference information may include a third vector, and the third vector includes at least one element. Each element in the at least one element corresponds to an AI model, and the value of each element indicates the degree of preference of the user for the AI model corresponding to the element. In some embodiments, the third preference information may indicate the AI model preferred by the user based on the service type corresponding to the AI model and the application to which the AI model belongs.
[0171] Alternatively, in some other embodiments, the third preference information may indicate the AI model preferred by the user based on the service type corresponding to the AI model.
[0172] S504. The terminal device fuses the second preference information and the third preference information to obtain first preference information.
[0173] By fusing the second preference information learned by the server and the third preference information learned by the terminal device to generate the first preference information, the problem of long-tail users of the model when the server learns the user's preference for the AI model can be improved, and the overfitting problem caused by less user data of the terminal device can be improved, thereby improving the accuracy of learning the user's preference for the AI model.
[0174] In some embodiments, the second preference information includes a second vector, and the third preference information includes a third vector. The terminal device may add the product of the second vector and a first coefficient and the product of the third vector and a second coefficient to obtain a first vector. The first vector is the first preference information. The first vector includes at least one element, and each element in the at least one element corresponds to an AI model, and the value of each element indicates the preference degree of the user for the AI model corresponding to the element.
[0175] Wherein, the sum of the first coefficient and the second coefficient is 1. It can be understood that the embodiments of the present application do not limit the manner of setting the first coefficient and the second coefficient and the specific magnitudes of the first coefficient and the second coefficient. Exemplarily, the first coefficient and the second coefficient may be set in advance by those skilled in the art, and both the first coefficient and the second coefficient are 0.5.
[0176] For example, if the second vector is [0.1, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.2], the third vector is [0.2, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.3], and both the first coefficient and the second coefficient are 0.5, then the first vector = the second vector * 0.5 + the third vector * 0.5, that is, the first vector is [0.15, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.25].
[0177] It can be understood that the manner of fusing the second preference information and the third preference information may correspond to the data types of the second preference information and the third preference information. When the second preference information and the third preference information are information of other data types, the terminal device may also adopt a fusion manner corresponding to the other data type to fuse the second preference information and the third preference information. Exemplarily, the data types of the second preference information and the third preference information are key-key value, and the terminal device may add the key values corresponding to the same key to obtain a new key value corresponding to the key.
[0178] In the embodiments of the present application, both the terminal device and the server can learn the preference of the terminal device for the AI model based on the user data of the terminal device, and respectively obtain the second preference information and the third preference information, and then fuse the second preference information and the third preference information to obtain the first preference information, which improves the problem of long-tail users of the model when the server learns the preference of the user for the AI model, and improves the overfitting problem caused by less user data of the terminal device, thereby being able to improve the accuracy of learning the preference of the user for the AI model.
[0179] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application provides a device for deploying an AI model. The device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this device embodiment. However, it should be clear that the device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.
[0180] In a possible implementation manner, the device for deploying an AI model may correspond to the terminal device in the foregoing method embodiment. For example, it may be a terminal device or a chip configured in the terminal device. The device for deploying an AI model is used to execute each step or process corresponding to the terminal device in the above method.
[0181] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application provides a device for deploying an AI model. The device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this device embodiment. However, it should be clear that the device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.
[0182] In a possible implementation manner, the device for deploying an AI model may correspond to the terminal device in the foregoing method embodiment. For example, it may be a server or a chip configured in the server. The device for deploying an AI model is used to execute each step or process corresponding to the server in the above method.
[0183] Based on the same inventive concept, an embodiment of the present application further provides a terminal device. The terminal device includes: a memory and a processor. The memory is used to store a computer program; the processor is used to execute the method described in the foregoing method embodiment when calling the computer program.
[0184] The terminal device provided in this embodiment can execute the foregoing method embodiment, and its implementation principle and technical effects are similar, so they will not be described here.
[0185] Based on the same inventive concept, an embodiment of the present application further provides a chip system. The chip system includes a processor, and the processor is coupled to a memory. The processor executes a computer program stored in the memory to implement the method described in the foregoing method embodiment.
[0186] Among them, the chip system may be a single chip or a chip module composed of multiple chips.
[0187] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing method embodiment is implemented.
[0188] The embodiments of the present application also provide a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the method described in the above method embodiments when executed.
[0189] Based on the same inventive concept, the embodiments of the present application also provide a server. The server includes: a memory and a processor, where the memory is used to store a computer program; the processor is used to execute the method described in the above method embodiments when calling the computer program.
[0190] The server provided in this embodiment can execute the above method embodiments, and its implementation principle and technical effects are similar, so they will not be elaborated here.
[0191] Based on the same inventive concept, the embodiments of the present application also provide a chip system. The chip system includes a processor, and the processor is coupled to a memory. The processor executes the computer program stored in the memory to implement the method described in the above method embodiments.
[0192] Among them, the chip system can be a single chip or a chip module composed of multiple chips.
[0193] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiments is implemented.
[0194] The embodiments of the present application also provide a computer program product. When the computer program product runs on a server, the server is enabled to execute the method described in the above method embodiments when executed.
[0195] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can at least include: any entity or device that can carry the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.
[0196] In the above embodiments, the descriptions of the respective embodiments each have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0197] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0198] In the embodiments provided in this application, it should be understood that the disclosed devices / apparatuses and methods can be implemented in other ways. For example, the device / apparatus embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0199] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0200] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0201] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0202] In addition, in the descriptions of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0203] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0204] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for deploying an artificial intelligence model, characterized in that: The method comprises: The terminal device sends hardware information, resource information and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; The server acquires, based on the hardware information, the resource information, and the first preference information, a first operator library matching the terminal device, where the first operator library includes at least one operator; The server returns the first operator library to the terminal device; The terminal device deploys the AI model based on the first operator library.
2. The method according to claim 1, characterized in that The method further comprises: The terminal device sends user data of the terminal device to the server; The server obtains second preference information based on the user data; The server returns the second preference information to the terminal device; The terminal device acquires third preference information based on the user data; The terminal device generates the first preference information based on the second preference information and the third preference information.
3. The method according to claim 1, characterized in that The method further comprises: The terminal device acquires third preference information based on user data of the terminal device; The terminal device sends the user data and the third preference information to the server; The server obtains second preference information based on the user data; The server generates the first preference information based on the second preference information and the third preference information.
4. The method according to claim 2 or 3, characterized in that: The server obtains second preference information based on the user data, including: The server inputs the user data into the first model to obtain the second preference information output by the first model; The terminal device acquires third preference information based on the user data, including: The terminal device inputs the user data into a second model to obtain the third preference information output by the second model, where the second model is obtained by compressing the first model.
5. The method according to claim 4, characterized in that The method further comprises: The server obtains the first model based on at least two user data; The server compresses the first model to obtain the second model; The server returns the second model to the terminal device.
6. The method according to any one of claims 2 to 5, characterized in that: The first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI model corresponding to the element.
7. The method according to claim 6, characterized in that The second preference information includes a second vector, the third preference information includes a third vector, and the terminal device generates the first preference information based on the second preference information and the third preference information, including: The terminal device adds the product of the second vector and the first coefficient and the product of the third vector and the second coefficient to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.
8. The method according to any one of claims 1 to 7, characterized in that: The terminal device deploys the AI model based on the first operator library, including: The terminal device compiles the AI model in the intermediate representation IR format based on the first operator library to obtain the AI model in the executable file format; and / or, The terminal device runs the AI model in the executable file format.
9. The method according to any one of claims 1 to 7, characterized in that: The terminal device further includes a preset second operator library, the second operator library includes at least one operator, and the terminal device deploys the AI model based on the first operator library, including: The terminal device deploys the AI model based on the first operator library and the second operator library.
10. A method for deploying an artificial intelligence model, characterized in that: Applied to a terminal device, the method comprises: Sending hardware information, resource information, and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; Acquire a first operator library returned by the server, where the first operator library includes at least one operator; The terminal device deploys the AI model based on the first operator library.
11. The method according to claim 10, characterized in that The method further comprises: Sending user data of the terminal device to the server; Acquire second preference information returned by the server based on the user data; acquiring third preference information based on the user data; The first preference information is generated based on the second preference information and the third preference information.
12. The method according to claim 11, characterized in that The first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI model corresponding to the element.
13. The method according to claim 12, characterized in that The second preference information includes a second vector, the third preference information includes a third vector, and generating the first preference information based on the second preference information and the third preference information includes: The product of the second vector and the first coefficient and the product of the third vector and the second coefficient are added to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.
14. A method for deploying an artificial intelligence model, characterized in that: Applied to a server, the method comprises: Acquire hardware information, resource information, and first preference information of a terminal device, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by a user of the terminal device; Based on the hardware information, the resource information, and the first preference information, acquiring a first operator library matching the terminal device, the first operator library including at least one operator; The first operator library is returned to the terminal device, where the first operator library is used for the terminal device to deploy the AI model.
15. The method according to claim 14, characterized in that The method further comprises: Acquiring user data and third preference information of the terminal device; acquiring second preference information based on the user data; The first preference information is generated based on the second preference information and the third preference information.
16. The method according to claim 15, characterized in that The method further comprises: acquiring a first model based on at least two user data, wherein the first model is used by the server to acquire the third preference information; compressing the first model to obtain a second model, where the second model is used by the terminal device to obtain the second preference information; The second model is returned to the terminal device.
17. A system, characterized in that: The system includes a server and a terminal device; The terminal device is used to send hardware information, resource information and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; deploying the AI model based on a first operator library, wherein the first operator library includes at least one operator; The server is used to obtain a first operator library matching the terminal device based on the hardware information, the resource information and the first preference information; and return the first operator library to the terminal device.
18. A terminal device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method according to any one of claims 10 to 13 when calling the computer program.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 10 to 13 is implemented.
20. A computer program product, characterized in that When the computer program product is executed on a terminal device, the terminal device is enabled to execute the method according to any one of claims 10 to 13.
21. A server, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method according to any one of claims 14 to 16 when calling the computer program.
22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 14 to 16 is implemented.
23. A computer program product, characterized in that When the computer program product is run on a server, the server is caused to execute the method according to any one of claims 14 to 16.