Method for deploying artificial intelligence model, terminal device, server and system

The terminal device sends hardware, resource and user preference information to the server, obtains matching operator databases and deploys AI models, solving the problem of limited performance of terminal device AI models and achieving efficient operator database matching and performance improvement.

WO2025113067A1PCT designated stage expired Publication Date: 2025-06-05HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/128536
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-10-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the prior art, the performance of the AI ​​model of the terminal device is limited by hardware performance, application scenarios and user preferences, resulting in low matching and efficiency of the operator library, which in turn affects the overall performance of the AI ​​model.

Method used

The terminal device sends hardware information, resource information and user preference information to the server. The server obtains the matching operator library based on this information and sends it to the terminal device. The terminal device deploys the AI ​​model based on the operator library.

Benefits of technology

The accuracy and matching of the operator library issued to the terminal device is realized, the performance of terminal device deployment AI model is improved, and the resources of terminal device and the resources of operation optimization tools are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128536_05062025_PF_FP_ABST
    Figure CN2024128536_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence, and provides a method for deploying an artificial intelligence model, a terminal device, a server and a system. The method comprises: a terminal device sends to a server hardware information, resource information and first preference information of the terminal device, the hardware information being used for indicating the hardware performance of the terminal device, the resource information being used for indicating the state of device resources of the terminal device, and the first preference information being used for indicating at least one artificial intelligence (AI) model preferred by a user of the terminal device; on the basis of the hardware information, the resource information and the first preference information, the server acquires a first operator library matched with the terminal device, the first operator library comprising at least one operator; the server returns the first operator library to the terminal device; and the terminal device deploys the AI model on the basis of the first operator library. The technical solution provided by the present application achieves the accuracy of issuing the operator library to the terminal device, and improves the performance of deploying the AI model by the terminal device, thus saving resources of the terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

Method, terminal device, server and system for deploying artificial intelligence model

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 1, 2023, with application number 202311658352.1 and application name “Method, terminal device, server and system for deploying artificial intelligence models”, and the Chinese patent application filed with the State Intellectual Property Office on January 8, 2024, with application number 202410028176.1 and application name “Method, terminal device, server and system for deploying artificial intelligence models”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI), and in particular to a method, terminal device, server, and system for deploying an AI model. Background Art

[0003] AI is a discipline that studies, develops, implements, and applies intelligence. It aims to simulate human cognitive abilities through machines, giving them a certain degree of human intelligence. For example, AI can be used in a variety of fields, including image recognition, natural language processing, speech processing, human-computer interaction, and artistic creation.

[0004] In existing technologies, servers can build AI models based on the hardware performance of common, popular terminal devices and deploy these AI models to multiple terminal devices, allowing them to run on each terminal device. However, due to the varying hardware performance, application scenarios, and user base of different terminal devices, the performance of AI models on terminal devices is very limited.

[0005] Summary of the Invention

[0006] In view of this, the present application provides a method, terminal device, server and system for deploying artificial intelligence models, which achieves the accuracy of sending operator libraries to terminal devices, improves the performance of deploying AI models on terminal devices, and saves resources of terminal devices.

[0007] In order to achieve the above-mentioned objectives, in a first aspect, an embodiment of the present application provides a method for deploying an artificial intelligence model, the method comprising: a terminal device sends hardware information, resource information and first preference information of the terminal device to a server, the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; the server obtains a first operator library matching the terminal device based on the hardware information, the resource information and the first preference information, the first operator library including at least one operator; the server returns the first operator library to the terminal device; the terminal device deploys the AI ​​model based on the first operator library.

[0008] In an embodiment of the present application, the server can obtain the first preference information, hardware information and resource information of the terminal device. The first preference information can indicate at least one AI model preferred by the user of the terminal device, the hardware information can indicate the hardware performance of the terminal device, and the resource information can indicate the status of the device resources of the terminal device. Therefore, based on the first preference information, hardware information and resource information, the first operator library is obtained from the server, so that the obtained first operator library can be matched with the hardware performance of the terminal device, the status of the device resources and the user's preference for the use of the AI ​​model of the terminal device, thereby achieving the accuracy of sending the operator library to the terminal device and improving the performance of the terminal device in deploying the AI ​​model. Since the terminal device does not need to deploy an optimization tool for searching operators and a large number of users use less operator libraries, the terminal device saves the device resources of the storage optimization tool and a large number of operator libraries, and also saves the resources of the terminal device running the optimization tool to search for operators.

[0009] In some embodiments, the hardware information may include at least one of the following information: the type of the terminal device, the version of the terminal device, the type of at least one hardware device in the terminal device (such as a chip, a processor, or a memory), the version of at least one hardware device in the terminal device, the platform of at least one hardware device in the terminal device, and the universality level of at least one hardware device.

[0010] In some embodiments, the resource information may include at least one of the computing speed of the processor of the terminal device, the power consumption of the terminal device, the power consumption of at least one hardware device in the terminal device, the remaining storage space size of the memory, and the like.

[0011] In some embodiments, the method also includes: the terminal device sends user data of the terminal device to the server; the server obtains second preference information based on the user data; the server returns the second preference information to the terminal device; the terminal device obtains third preference information based on the user data; the terminal device generates the first preference information based on the second preference information and the third preference information.

[0012] In some embodiments, user data may include a historical record of a user using at least one AI model and / or a historical record of a user using at least one application. Based on the historical record of a user using at least one AI model and / or the historical record of a user using at least one application, information such as the frequency and duration of the user's use of at least one AI model can be determined, and the user's preference for at least one AI model can be obtained, that is, the first preference information can be obtained.

[0013] In some embodiments, the method further includes: the terminal device obtains third preference information based on user data of the terminal device; the terminal device sends the user data and the third preference information to the server; the server obtains second preference information based on the user data; and the server generates the first preference information based on the second preference information and the third preference information.

[0014] Both the terminal device and the server can learn the terminal device's preference for the AI ​​model based on the user data of the terminal device, and obtain second preference information and third preference information respectively, and then fuse the second preference information and the third preference information to obtain the first preference information. This improves the long-tail user problem of the model when the server learns the user's preference for the AI ​​model, and improves the overfitting problem of the terminal device caused by less user data, thereby improving the accuracy of learning the user's preference for the AI ​​model.

[0015] In some embodiments, the server may include at least one optimization tool, through which a first operator library matching the terminal device is obtained based on the hardware information, the resource information, and the first preference information. In some embodiments, the optimization tool may include a TBE search tool, a cube computing engine (CCE) search tool, a CPU search tool, a heterogeneous search tool, a function flow template library (FFTL) search tool, or a quantization tool. The TBE search tool can be used to generate constant NPU operators. The CCE search tool can be used to search the output knowledge base and generate constant NPU operators based on CCE general operators. The CPU search tool can be used to search for constant CPU operators. The heterogeneous search tool can generate NPU operators and CPU operators based on a preset heterogeneous strategy. In some embodiments, in some embodiments, the first operator library may include at least one operator among TBE operators generated by a TBE search tool, CCE operators and / or NPU operators generated by a CCE search tool, CPU operators generated by a CPU search tool, NPU operators and / or CPU operators generated by a heterogeneous search tool, operators generated by an FFTL search tool, and operators generated by a quantization tool.

[0016] In some embodiments, the server obtains second preference information based on the user data, including: the server inputs the user data into a first model, and obtains the second preference information output by the first model; the terminal device obtains third preference information based on the user data, including: the terminal device inputs the user data into a second model, and obtains the third preference information output by the second model, where the second model is obtained by compressing the first model.

[0017] In some embodiments, the structure of the second model is simpler than that of the first model, and / or the operating efficiency of the second model is higher than that of the first model, and / or the parameters included in the second model are fewer than those included in the first model. Therefore, the first model can be called a personalized large model and the second model can be called a personalized small model.

[0018] In some embodiments, the method further includes: the server acquiring the first model based on at least two user data;

[0019] The server compresses the first model to obtain the second model; and the server returns the second model to the terminal device.

[0020] In some embodiments, the compression process for the first model includes distillation, pruning, or quantization. Distillation involves using a larger first model to guide the training of a smaller second model, so that the performance of the trained second model is close to that of the first model. Pruning involves trimming and adjusting less important network structures in the first model to obtain a second model with a simpler network structure. Quantization involves reducing the computational precision of the data in the first model, thereby obtaining a faster-running second model.

[0021] In some embodiments, the first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

[0022] In some embodiments, the second preference information includes a second vector, the third preference information includes a third vector, and the terminal device generates the first preference information based on the second preference information and the third preference information, including: the terminal device adds the product of the second vector and the first coefficient and the product of the third vector and the second coefficient to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.

[0023] In some embodiments, the terminal device deploys the AI ​​model based on the first operator library, including: the terminal device compiles the AI ​​model in the intermediate representation IR format based on the first operator library to obtain the AI ​​model in the executable file format; and / or the terminal device runs the AI ​​model in the executable file format.

[0024] In some embodiments, the terminal device may compile an AI model in an executable file format based on at least one operator in the first operator library, thereby enabling, in the compiled state, the loading of optimization parameters that match the user preferences of the terminal device, the hardware performance of the device, and the status of the device resources, such as quantization weights, fusion operators, and overall fusion parameters. In some embodiments, the terminal device may load and run the AI ​​model in an executable file format based on at least one operator in the first operator library, thereby enabling the running of the AI ​​model based on optimization parameters that match the user preferences of the terminal device, the hardware performance of the device, and the status of the device resources.

[0025] In some embodiments, the terminal device also includes a pre-set second operator library, the second operator library includes at least one operator, and the terminal device deploys the AI ​​model based on the first operator library, including: the terminal device deploys the AI ​​model based on the first operator library and the second operator library.

[0026] In second aspect, an embodiment of the present application provides a method for deploying an artificial intelligence model, which is applied to a terminal device, and the method includes: sending hardware information, resource information and first preference information of the terminal device to a server, the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; obtaining a first operator library returned by the server, the first operator library including at least one operator; the terminal device deploys the AI ​​model based on the first operator library.

[0027] In some embodiments, the method further includes: sending user data of the terminal device to the server; obtaining second preference information returned by the server based on the user data; obtaining third preference information based on the user data; and generating the first preference information based on the second preference information and the third preference information.

[0028] In some embodiments, the first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

[0029] In some embodiments, the second preference information includes a second vector, the third preference information includes a third vector, and generating the first preference information based on the second preference information and the third preference information includes: adding the product of the second vector and the first coefficient and the product of the third vector and the second coefficient to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.

[0030] In a third aspect, an embodiment of the present application provides a method for deploying an artificial intelligence model, which is applied to a server, and the method includes: obtaining hardware information, resource information and first preference information of a terminal device, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; based on the hardware information, the resource information and the first preference information, obtaining a first operator library matching the terminal device, the first operator library including at least one operator; returning the first operator library to the terminal device, and the first operator library is used to deploy the AI ​​model on the terminal device.

[0031] In some embodiments, the method further includes: obtaining user data and third preference information of the terminal device; obtaining second preference information based on the user data; and generating the first preference information based on the second preference information and the third preference information.

[0032] In some embodiments, the method further includes: obtaining a first model based on at least two user data, the first model being used by the server to obtain the third preference information; compressing the first model to obtain a second model, the second model being used by the terminal device to obtain the second preference information; and returning the second model to the terminal device.

[0033] In a fourth aspect, an embodiment of the present application provides a system comprising a server and a terminal device; the terminal device is configured to send hardware information, resource information, and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence (AI) model preferred by the user of the terminal device; the AI ​​model is deployed based on a first operator library, the first operator library including at least one operator; the server is configured to obtain a first operator library matching the terminal device based on the hardware information, the resource information, and the first preference information; and return the first operator library to the terminal device.

[0034] In a fifth aspect, an embodiment of the present application provides a device for deploying an artificial intelligence model, which has the function of implementing the behavior of the terminal device in the above aspects and possible implementation methods of the above aspects. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a transceiver module or unit, a processing module or unit, an acquisition module or unit, etc.

[0035] In a sixth aspect, an embodiment of the present application provides a device for deploying an artificial intelligence model, which has the function of implementing the above-mentioned aspects and the possible implementation methods of the server. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above-mentioned functions. For example, a transceiver module or unit, a processing module or unit, an acquisition module or unit, etc.

[0036] In the seventh aspect, an embodiment of the present application provides a terminal device, comprising: a memory and a processor, the memory being used to store a computer program; the processor being used to execute any one of the methods described in the above second aspect when calling the computer program.

[0037] In an eighth aspect, an embodiment of the present application provides a chip system, which includes a processor coupled to a memory, and the processor executes a computer program stored in the memory to implement any method described in the second aspect above.

[0038] The chip system may be a single chip or a chip module composed of multiple chips.

[0039] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements any of the methods described in the second aspect when the computer program is executed by a processor.

[0040] In a tenth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute any of the methods described in the second aspect above.

[0041] In the eleventh aspect, an embodiment of the present application provides a server, comprising: a memory and a processor, the memory being used to store a computer program; the processor being used to execute any one of the methods described in the third aspect above when calling the computer program.

[0042] In the twelfth aspect, an embodiment of the present application provides a chip system, which includes a processor, the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement any one of the methods described in the third aspect above.

[0043] The chip system may be a single chip or a chip module composed of multiple chips.

[0044] In a thirteenth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above-mentioned third aspects is implemented.

[0045] In a fourteenth aspect, an embodiment of the present application provides a computer program product, which, when running on a server, enables the server to execute any of the methods described in the third aspect above.

[0046] It can be understood that the beneficial effects of the second to fourteenth aspects can be found in the relevant description of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG1 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0048] FIG2 is a structural block diagram of a system provided in an embodiment of the present application;

[0049] FIG3 is a flow chart of a method for deploying an AI model provided in an embodiment of the present application;

[0050] FIG4 is a schematic diagram of an AI model deployment process provided in an embodiment of the present application;

[0051] FIG5 is a flow chart of a method for learning user preference information provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The method for deploying a machine learning model provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0053] 1 is a schematic diagram of the structure of an electronic device 100 according to an embodiment of the present application. The electronic device 100 may include a processor 110, a memory 120, a communication module 130, and the like.

[0054] Among them, the processor 110 may include one or more processing units, and the memory 120 is used to store program code and data. In an embodiment of the present application, the processor 110 can execute computer-executable instructions stored in the memory 120 to control and manage the actions of the electronic device 100. For example, the processor 110 may include a system-on-a-chip (SOC). For example, the processor 110 may include a graphics processing unit (GPU), a neural-network processing unit (NPU) or a central processing unit (CPU). In some embodiments, the processor 110 may include an AI core and an AI processor. The AI ​​core can be used to run operators of specific algorithms, such as matrix, vector, and scalar calculation-intensive operators. The AI ​​processor can be used as a supplement to the AI ​​processor to run other operators.

[0055] The communication module 130 can be used for communication between the various internal modules of the electronic device 100, or for communication between the electronic device 100 and other external electronic devices. For example, if the electronic device 100 communicates with other electronic devices via a wired connection, the communication module 130 may include an interface, such as a USB interface. The USB interface may be an interface that complies with USB standards, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other electronic devices.

[0056] Alternatively, the communication module 130 may include an audio device, a radio frequency circuit, a Bluetooth chip, a wireless fidelity (Wi-Fi) chip, a near-field communication (NFC) module, etc., and can realize the interaction between the electronic device 100 and other electronic devices in a variety of different ways.

[0057] The wireless communication function of the electronic device 100 can be implemented through at least one antenna, a mobile communication module, a wireless communication module, a modem processor, and a baseband processor.

[0058] The antenna is used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, an antenna can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antenna can be used in conjunction with a tuning switch.

[0059] Optionally, the electronic device 100 may further include a display screen 140, which may display images or videos in a human-computer interaction interface. The electronic device 100 may implement a display function through an image processing microprocessor, a display screen, and an application processor.

[0060] Display screen 140 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0061] The electronic device 100 can implement a shooting function through a camera, a video codec, a GPU, a display screen 140, and an application processor.

[0062] Optionally, the electronic device 100 may further include a peripheral device 150 , such as a mouse, a keyboard, a speaker, a microphone, and the like.

[0063] Optionally, the electronic device 100 may further include a charging management module, a power management module, a battery, a button, an indicator, and one or more SIM card interfaces, etc., and this embodiment of the application does not impose any restrictions on this.

[0064] It should be understood that, other than the various components or modules listed in FIG1 , the embodiments of the present application do not specifically limit the structure of the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0065] First, some of the terms involved in the embodiments of this application are explained:

[0066] AI is a discipline that studies, develops, implements, and applies intelligence. It aims to simulate human cognitive abilities through machines, enabling them to possess a certain degree of human intelligence. For example, AI can be used to achieve capabilities in a variety of fields, including image recognition, natural language processing, speech processing, human-computer interaction, and artistic creation.

[0067] An AI model is a mathematical model that analyzes, processes, predicts, and optimizes data with certain regularity and predictability. A corresponding AI model can be built based on the needs of a specific business or capability field, and the AI ​​model has the ability to handle that specific business. For example, according to the business type of the AI ​​model, the optical character recognition (OCR) model can recognize text in an image, the natural language understanding (NLU) model can understand the semantics of human language and recognize user intent, the natural language generation (NLG) model can generate text information that humans can understand, the text-to-speech (TTS) model can generate voice information based on text information, the face recognition model can recognize faces included in an image, and the super-resolution model can increase the resolution of the input image to achieve high definition.

[0068] Exemplarily, AI models may include linear models, decision tree models, ensemble models, neural network models, and support vector machine models. A linear model is a simple model that can divide or predict data using a straight line or hyperplane. Common linear models include linear regression, logistic regression, etc. A decision tree model is a model based on a tree structure that divides a data set into smaller subsets until each subset contains data points of a single category, such as classification tree models and regression tree models. An ensemble model is a model that improves prediction performance by combining multiple models, such as random forest models and gradient boosting tree models. A neural network model is a model based on the biological nervous system that establishes complex mapping relationships through connections between multiple neurons, such as multi-layer perceptron models, convolutional neural network models, and recurrent neural network models. A support vector machine model is a model based on maximum margin classification that maps data points into a high-dimensional space and finds the maximum margin hyperplane to divide the data, such as linear support vector machine models and nonlinear support vector machine models.

[0069] In some embodiments, an AI model may include a computational graph and weights.

[0070] A computation graph is a formalized method for representing the operation of an algorithm. A computation graph consists of multiple nodes connected by directed edges, with each node representing an operator. The input edge entering a node represents the input data of the operator corresponding to that node, and the output edge leaving a node represents the output data of the operator corresponding to that node. The operation process represented by the computation graph can be the model inference process or the model training process. By saving each operator and its associated weights according to the connection method of the directed edges in the computation graph, the AI ​​model can be saved.

[0071] Operators are the basic components of AI models and can implement one or more operations in AI models. For example, the convolution operator can be used to extract and map features of data; the activation function operator can be used to add nonlinear factors to improve the expressiveness of the AI ​​model and solve problems that linear models cannot solve; the loss function operator can be used to measure the error between the predicted results and the actual results of the AI ​​model, and then train, evaluate or optimize the AI ​​model based on the error. In some embodiments, the operator in the AI ​​model may correspond to a layer in the AI ​​model. In some embodiments, the parameters of the operator may include the input data of the operator, the output data of the operator, and the elements in the operator (such as the size of the convolution kernel in the convolution operator).

[0072] Weights are used to represent the data that operators need to use during execution.

[0073] Tasks are smaller scheduling units derived from operator decomposition. They enable inter-operator and intra-operator scheduling, breaking operator boundaries and allowing fine-grained scheduling of computations to hardware devices. In some implementations, an operator can be divided into multiple task units.

[0074] An operator core (kernel) is a specific implementation of an operator based on the hardware device that executes the operator. The operator core can be obtained by compiling the operator. The operator core implements the computing logic of a specific task unit and determines the total number of task units. In some embodiments, when the same operator is compiled for different hardware devices (such as CPU and NPU), the resulting operator cores are also different. In some embodiments, an operator can be implemented as one or more operator cores. In some embodiments, the operator core can include files in *.so format and files in *.json format.

[0075] The model formats of AI models can include the following formats.

[0076] A framework model is an AI model in the model format output by a deep learning framework. For example, a deep learning framework may include a Hopfield neural network (HNN) or a mind spore. The framework model is independent of the hardware that actually runs the AI ​​model.

[0077] The intermediate representation (IR) model can be used to represent the structure and computational process of an AI model in the form of a computational graph. The IR model is independent of the hardware that actually runs the AI ​​model.

[0078] An AI model in executable file format. An executable file is a file that can be run on hardware. The executable file format is related to the hardware that actually runs the AI ​​model. Different types of hardware can correspond to different executable files. For example, if the hardware running the AI ​​model on the terminal device is a CPU, the executable file can correspond to that CPU.

[0079] In some implementations, since the framework model and IR model are independent of the hardware on which the AI ​​model actually runs, and the executable file format is related to the hardware on which the AI ​​model actually runs, the framework model and IR model can also be in a high-level model format, and the executable file format can be in a low-level model format. The AI ​​model in the high-level model format can be compiled for the hardware on which the AI ​​model runs using a compilation tool to obtain an AI model in a low-level model format.

[0080] In some embodiments, the compilation tool may include a tensor virtual machine (TVM) or a tensor boost engine (TBE). It is understood that in actual applications, the AI ​​model can also be compiled using other compilation tools.

[0081] Taking TVM as an example, the compilation process of the AI ​​model by the compilation tool may include: obtaining the AI ​​model in a high-level model format and generating a computational graph for the AI ​​model; optimizing the computational graph (for example, rewriting high-level data flows and / or optimizing operator levels) to obtain an optimized computational graph; and compiling and generating an executable file that can be run by hardware devices such as processors in terminal devices based on the optimized operator graph.

[0082] In order to facilitate understanding of the technical solutions in the embodiments of the present application, the application scenarios of the embodiments of the present application are introduced below.

[0083] To achieve intelligent terminal devices, various AI models can be deployed in these devices, each capable of handling at least one service. When deploying AI models, improving the performance experience of the services handled by these AI models is a critical issue. For example, for terminal devices with poor hardware performance, deploying and optimizing AI models on these devices can improve the performance of the AI ​​models in handling related services on the terminal devices, significantly reducing the performance gap between the service experience on these devices and that of other terminals with better hardware performance.

[0084] As can be seen from the foregoing, operators are the fundamental components of AI models. Therefore, based on the hardware information of the terminal device, an optimization tool can be used to search for an operator library for the AI ​​model. This operator library includes at least one operator and can indicate the model structure and operating parameters of the AI ​​model. The terminal device deploys the AI ​​model based on the searched operator library, ensuring that the deployed AI model has a better model structure and model operation process.

[0085] In some embodiments, the optimization tool may include at least one of a TBE search tool, a CCE search tool, a CPU search tool, a heterogeneous search tool, an FFTL search tool, and a quantization tool. The TBE search tool may be used to generate constant NPU operators. The CCE search tool may be used to search an output knowledge base and generate constant NPU operators based on CCE general operators. The CPU search tool may be used to search for constant CPU operators. The heterogeneous search tool may generate NPU operators and CPU operators based on a preset heterogeneous strategy.

[0086] In some embodiments, the server includes an optimization tool. Based on the hardware performance of common, popular terminal devices, the server can use this optimization tool to search for a universal operator library and send this operator library to each terminal device, allowing each terminal device to deploy an AI model based on this operator library. However, different terminal devices have different product forms, product versions, application scenarios, and application preferences. Therefore, this method of uniformly sending a universal operator library to terminal devices cannot produce an operator library suitable for each terminal device, resulting in poor performance of the deployed AI model.

[0087] In some embodiments, a terminal device includes an optimization tool. The terminal device can search for an operator library that matches the terminal device based on the terminal device's hardware information. However, since the optimization tool requires memory on the terminal device, and searching for the optimal operator using the optimization tool also requires memory and power consumption, terminal devices with limited resources such as read-only memory (ROM) and random access memory (RAM) cannot store and apply the optimization tool.

[0088] To address at least some of the above technical issues, an embodiment of the present application provides another method for deploying AI models. The server can use optimization tools to search for an operator library that matches the terminal device based on the terminal device's hardware information, user preferences, and resource information. The terminal device can dynamically manage the required operator library based on user usage habits and frequency, minimizing the space required by the terminal device. During inference, the terminal device achieves performance optimization during compilation and runtime by deploying the operator library that matches the terminal device online.

[0089] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0090] Please refer to Figure 2, which is a schematic diagram of the structure of a system 200 provided in an embodiment of the present application. System 200 may include a server 210 and a terminal device 220. Server 210 and terminal device 220 may be connected via a network. Server 210 and terminal device 220 may be implemented using the electronic device 100 described above.

[0091] The server 210 may include an optimization tool. Using the optimization tool, the server 210 may search for a dedicated operator library that matches the terminal device 220 based on relevant information about the terminal device 220, such as hardware information, user preferences, and resource information. The dedicated operator library may include at least one operator that matches the terminal device 220. The server 210 may publish the dedicated operator library so that the terminal device 220 can obtain the dedicated operator library.

[0092] The terminal device 220 includes an application program, an AI operator library, and a processor.

[0093] Applications can include or call AI models. The terminal device 220 can include multiple applications. The same application can include or call multiple AI models of different business types. Different applications can also include or call AI models of the same business type. For example, video application A can include a super-resolution model, video application B can also include a super-resolution model, and a voice assistant can include an automatic speech recognition (ASR) model and an OCR model.

[0094] The AI ​​operator library may include a dedicated operator library that matches the terminal device 220. In some embodiments, the terminal device 220 may dynamically download the dedicated operator library from the server 210. In some embodiments, the terminal device 220 may also include a general operator library, which may include at least one operator. The general operator library may be the same for different terminal devices 220.

[0095] The processor can be used to run one or more operators to implement the business processed by applications and AI models.

[0096] Based on the system 200 shown in Figure 2, the server 210 can search for and publish a dedicated operator library that matches the terminal device 220 based on the hardware information of the terminal device 220, the user's preference information, resource information, and other related information. The terminal device 220 can download the dedicated operator library from the server 210 and deploy an AI model based on the operators in the dedicated operator library, thereby achieving performance optimization during compilation and runtime.

[0097] Please refer to Figure 3, which is a flowchart of a method for deploying an AI model provided in an embodiment of the present application. Among them, the method can be applied to the system shown in Figure 2. It should be noted that the method is not limited to Figure 3 and the specific order described below. It should be understood that in other embodiments, the order of some steps in the method can be interchanged according to actual needs, or some steps can be omitted or deleted. The method includes the following steps:

[0098] S301: The terminal device and / or the server performs personalized learning on the user preference of the terminal device to obtain first preference information.

[0099] The first preference information may be used to indicate at least one AI model preferred by a user of the terminal device. In some embodiments, the first preference information may also be used to indicate a degree of preference of the user of the terminal device for each of the at least one AI model.

[0100] In some embodiments, the terminal device and / or the server may determine the first preference information based on the user data of the terminal device. The user data may be obtained by the terminal device during the user's use of the terminal device. In some embodiments, the user data may include a history of the user's use of at least one AI model and / or a history of the user's use of at least one application. Based on the history of the user's use of at least one AI model and / or the history of the user's use of at least one application, information such as the frequency and duration of the user's use of the at least one AI model may be determined, and then the user's preference for the at least one AI model may be obtained, that is, the first preference information may be obtained. The user's preference for the AI ​​model may be positively correlated with the frequency and duration of the user's use of the AI ​​model.

[0101] In some embodiments, a terminal device sends user data of the terminal device to a server, the server obtains second preference information based on the user data, the server returns the second preference information to the terminal device, the terminal device obtains third preference information based on the user data, and the terminal device generates first preference information based on the second and third preference information. The second preference information may indicate at least one AI model preferred by the user learned by the server, and the third preference information may indicate at least one AI model preferred by the user learned by the terminal device. That is, the first preference information can be jointly determined by the second and third preference information learned by the terminal device and the server based on the user data of the terminal device, thereby improving the accuracy of determining the first preference information.

[0102] In some embodiments, the server includes a first model, and the server can input user data of the terminal device into the first model to obtain second preference information output by the first model.

[0103] The first model can receive input user data and output corresponding preference information. In some embodiments, the server can obtain user data from at least two terminal devices and train the first model based on the obtained at least two user data. Alternatively, in other embodiments, the server can obtain a pre-trained first model. In some embodiments, if the server obtains new user data, it can also perform incremental training on the first model based on the new user data, thereby improving the accuracy of the first model.

[0104] In some embodiments, the terminal device includes a second model, and the terminal device can input user data of the terminal device into the second model to obtain third preference information output by the second model.

[0105] The second model can receive input user data and output corresponding preference information. In some embodiments, the terminal device can train and obtain the second model based on the user data of the terminal device. Alternatively, in other embodiments, the terminal device can obtain a trained second model. For example, the terminal device can obtain the second model from a server. In some embodiments, if the terminal device obtains new user data, the second model can be incrementally trained based on the new user data, thereby improving the accuracy of the second model.

[0106] It is understandable that the first model and the second model can be the same or similar, and the way the server obtains the first model can be the same or similar to the way the terminal device obtains the second model. In some embodiments, the structure of the second model is simpler than that of the first model, and / or the operating efficiency of the second model is higher than that of the first model, and / or the parameters included in the second model are fewer than those included in the first model. Therefore, the first model can be called a personalized large model, and the second model can be called a personalized small model.

[0107] In some implementations, the terminal device acquires the first preference information based on user data of the terminal device, that is, the first preference information is determined only by the terminal device.

[0108] In some embodiments, the terminal device may receive first preference information submitted by a user. For example, the terminal device may display an interface for receiving the first preference information to the user and receive the first preference information submitted by the user through the interface. Alternatively, the terminal device may receive a file containing the first preference information submitted by the user and obtain the first preference information from the file.

[0109] In some embodiments, the terminal device obtains third preference information based on user data of the terminal device, the terminal device sends the user data and the third preference information to the server, the server obtains second preference information based on the user data, and the server generates first preference information based on the second preference information and the third preference information.

[0110] In some implementations, the terminal device sends user data of the terminal device to the server, and the server obtains the first preference information based on the user data, that is, the first preference information is determined only by the server.

[0111] In addition, the manner in which the terminal device and the server perform personalized learning on the user preference of the terminal device to obtain the first preference information can also be referred to the detailed description in FIG. 5 below.

[0112] S302: The terminal device obtains a first operator library from the server based on the first preference information, hardware information, and resource information.

[0113] The first preference information can indicate at least one AI model preferred by the user of the terminal device, the hardware information can indicate the hardware performance of the terminal device, and the resource information can indicate the status of the device resources of the terminal device. Therefore, based on the first preference information, hardware information, and resource information, the first operator library is obtained from the server, and the obtained first operator library can be matched with the hardware performance of the terminal device, the status of the device resources, and the user's preference for using AI models of the terminal device, thereby achieving the accuracy of sending the operator library to the terminal device. Since the terminal device does not need to deploy optimization tools for searching operators and a large number of users use less operator libraries, the terminal device saves the device resources of storing optimization tools and a large number of operator libraries, and also saves the resources of the terminal device running optimization tools to search for operators.

[0114] In some embodiments, the hardware information may include at least one of the following information: the type of the terminal device, the version of the terminal device, the type of at least one hardware device in the terminal device (such as a chip, a processor, or a memory), the version of at least one hardware device in the terminal device, the platform of at least one hardware device in the terminal device, and the universality level of at least one hardware device.

[0115] In some embodiments, the resource information may include at least one of the computing speed of the processor of the terminal device, the power consumption of the terminal device, the power consumption of at least one hardware device in the terminal device, the remaining storage space size of the memory, and the like.

[0116] In some implementations, the terminal device may send first preference information, hardware information, and resource information to the server. The server, based on the first preference information, hardware information, and resource information, searches for a first operator library that matches the terminal device and sends the first operator library to the terminal device. It is understood that if the first preference information is determined by the server, the terminal device may send hardware information and resource information to the server.

[0117] In some embodiments, the server may search for a first operator library that matches the first preference information, hardware information, and resource information from a plurality of stored operator libraries based on at least one optimization tool. In some embodiments, the first operator library may include a TBE operator generated by a TBE search tool, a CCE operator and / or an NPU operator generated by a CCE search tool, a CPU operator generated by a CPU search tool, an NPU operator and / or a CPU operator generated by a heterogeneous search tool, an operator generated by an FFTL search tool, and at least one operator generated by an operator generated by a quantization tool. The server may obtain and store a plurality of operator libraries in advance. Exemplarily, the plurality of operator libraries may include a general operator library, a CV class operator library, an L0-level device operator library, and an operator library that uses only a video model user. It is understandable that the embodiment of the present application does not limit the specific contents of the plurality of operator libraries included in the server.

[0118] S303: The terminal device deploys the AI ​​model based on the first operator library.

[0119] When the terminal device obtains the first operator library that matches the first preference information, hardware information and resource information of the terminal device from the server, the AI ​​model can be deployed based on the first operator library, thereby achieving optimized deployment of the AI ​​model based on the user preferences of the terminal device, the hardware performance of the device and the status of the device resources.

[0120] The AI ​​model may be an AI model to be deployed. In some implementations, when the terminal device receives a request from an application to deploy an AI model, it deploys the AI ​​model requested by the application based on the first operator library.

[0121] In some embodiments, the terminal device compiles the IR model (i.e., the AI ​​model in IR format) based on the first operator library to obtain the AI ​​model in executable file format; and / or the terminal device can run the AI ​​model in executable file format. In some embodiments, the terminal device can compile the AI ​​model in executable file format based on at least one operator in the first operator library, so as to achieve, in the compiled state, mounting optimization parameters that match the user preferences of the terminal device, the hardware performance of the device, and the status of the device resources, such as quantization weights, fusion operators, overall fusion parameters, etc. In some embodiments, the terminal device can load and run the AI ​​model in executable file format based on at least one operator in the first operator library, so as to achieve running the AI ​​model based on optimization parameters that match the user preferences of the terminal device, the hardware performance of the device, and the status of the device resources.

[0122] In some embodiments, the terminal device can compile the operators in the AI ​​model into one or more operator cores and one or more task units based on the hardware device in the terminal device for executing the operators, and generate code and instructions based on the one or more operator cores and one or more task units to obtain the AI ​​model in an executable file format.

[0123] Exemplarily, the manner in which the terminal device deploys the AI ​​model based on the first operator library may be as shown in FIG4 . The first terminal 230 may be a terminal device held by a user.

[0124] Server 210 includes an optimization tool that searches for and obtains a first operator library that matches the user preferences, device hardware performance, and device resource status of first terminal 230. Server 210 publishes the first operator library. In some embodiments, the optimization tool can perform operator search and compilation, model conversion, record weight reordering rules, and record quantization parameters to obtain the first operator library.

[0125] The first terminal 230 obtains and stores the first operator library from the server 210. The first terminal 230 can obtain a framework model through a deep learning framework, convert the framework model into an IR model, and compile the IR model based on the first operator library, thereby mounting the task units, operator cores, and weight reordering rules carried by the first operator library to obtain an AI model in an executable file format. The first terminal 230 can store the AI ​​model in an executable file format, as well as load and run the AI ​​model in an executable file format.

[0126] In some embodiments, before publishing the first operator library, the server 210 may also send the first operator library to the second terminal 240 so that relevant technical personnel can test the first operator library through the second terminal 240. Deploying an AI model based on the first operator library can determine the performance improvement of the deployed AI model by the first operator library. When it is determined that the first operator library can improve the performance of the deployed AI model, the first operator library can be published on the server 210. The second terminal 240 can be a test machine held by relevant technical personnel, and the second terminal 240 can be similar to or the same as the first terminal device 230.

[0127] In some implementations, a partial file structure of an AI model in an executable file format may be as follows:

[0128] / data / hiai / hash_id / #hash_id is the hash value of the calculation graph output by the IR model

[0129] |----main /

[0130] |----model_info.json #Configuration file of AI model

[0131] |----xyz.omc #AI model excluding weights

[0132] |----subgraph_0 / #TVM operator

[0133] |----xxx.so #library file

[0134] |----yyy.json #weight conversion record file

[0135] |----subgraph_1 / #TBE operator

[0136] |----zzz.json #weight conversion record file

[0137] |----subgraph_2 / #CPU computing language (CPUCL) operator

[0138] |----aaa.json #weight conversion record file

[0139] |----bbb.so #library file, including the CPUCL operator library operator core used

[0140] |----subgraph_3 / #CCE operator

[0141] |----ccc.json #weight conversion record file

[0142] Among them, subgraph is the weight conversion configuration file of the sub-computation graph. Each AI model can include multiple subgraphs, and each subgraph corresponds to the weight conversion configuration file of the sub-computation graph of a type of operator.

[0143] In some embodiments, the configuration file of the AI ​​model may include at least one of the following information: the file name, version number, and segmentation information of the sub-computation graph of the AI ​​model.

[0144] In some embodiments, the terminal device further includes a pre-installed second operator library, which may include at least one operator. The terminal device may deploy an AI model based on the first and second operator libraries. In some embodiments, the operators in the second operator library may be general-purpose operators.

[0145] In some implementations, the differences between the first operator library and the second operator library may include:

[0146] The first difference is that the first operator library includes dedicated operators (also called specific operators) deployed based on the first preference information of the terminal device, the device's hardware information, and resource information, while the second operator library includes general operators. In some embodiments, the first operator library can be the customized operator library in Figure 2, and the second operator library can be the general operator library in Figure 2.

[0147] The second difference is that the first operator library is dynamically deployed based on the dynamic changes in the first preference information, hardware information, and resource information of the terminal device. When at least one of the first preference information, hardware information, or resource information changes, the operators included in the first operator library are also updated. The second operator library is not updated based on changes in the first preference information, hardware information, or resource information.

[0148] In an embodiment of the present application, the server can obtain the first preference information, hardware information and resource information of the terminal device. The first preference information can indicate at least one AI model preferred by the user of the terminal device, the hardware information can indicate the hardware performance of the terminal device, and the resource information can indicate the status of the device resources of the terminal device. Therefore, based on the first preference information, hardware information and resource information, the first operator library is obtained from the server, so that the obtained first operator library can be matched with the hardware performance of the terminal device, the status of the device resources and the user's preference for the use of the AI ​​model of the terminal device, thereby achieving the accuracy of sending the operator library to the terminal device and improving the performance of the terminal device in deploying the AI ​​model. Since the terminal device does not need to deploy an optimization tool for searching operators and a large number of users use less operator libraries, the terminal device saves the device resources of the storage optimization tool and a large number of operator libraries, and also saves the resources of the terminal device running the optimization tool to search for operators.

[0149] Please refer to Figure 5, which is a flow chart of a method for obtaining preference information provided in an embodiment of the present application. The method may be a detailed description of S301 described above. It should be noted that the method is not limited to the specific sequence described in Figure 5 and below. It should be understood that in other embodiments, the order of some steps in the method may be interchanged according to actual needs, or some steps may be omitted or deleted. The method includes the following steps:

[0150] S501: The server trains a first model based on user data of at least two terminal devices.

[0151] The user data of the at least two terminal devices may include historical records of the use of the AI ​​model by the users of the at least two terminal devices and / or records of at least one application by the users of the at least two terminal devices. The server trains and obtains the first model based on the user data of the at least two terminal devices, that is, learns the preferences of the group users for the AI ​​model.

[0152] In some embodiments, the server may receive user data sent by at least two terminal devices respectively. In some embodiments, the user data may be non-private data.

[0153] In some embodiments, the server may train the first model at intervals of a first duration or upon obtaining a first amount of user data, thereby ensuring the accuracy of the first model. It should be noted that the embodiments of the present application do not limit the manner in which the first duration and the first number are set, nor do they limit the specific sizes of the first duration and the first number. For example, the first duration and the first number may be set or updated in advance by relevant technical personnel.

[0154] In some implementations, the server may obtain the first model through other means. For example, the server may obtain a first model that has already been trained, so S501 may be omitted.

[0155] S502: The server compresses the first model to obtain a second model.

[0156] Compressing the first model, i.e., miniaturizing it, can make the resulting second model run faster and consume fewer device resources. Because terminal devices have fewer device resources than servers, the server can compress the first model to produce a second model that is easier to deploy on the terminal device, allowing the terminal device to learn the user's preferences for AI models from user data.

[0157] In some embodiments, the compression processing performed on the first model includes distillation, pruning, or quantization.

[0158] Distillation is the process of using a larger first model to guide the training of a smaller second model. The effect of the trained second model can be close to that of the first model.

[0159] Pruning is to trim the less important network structures in the first model to obtain a second model with a simpler network structure.

[0160] Quantization is the process of reducing the computational precision of the data in the first model, thereby obtaining a second model that runs faster.

[0161] It is understandable that, in practical applications, the first model may also be compressed in other ways.

[0162] In some implementations, if the server updates the first model, the server may update the second model deployed on the terminal device based on the updated first model.

[0163] In some implementations, the terminal device may obtain the second model through other means. For example, the terminal device may obtain the second model through training based on user data of the terminal device, so S502 may be omitted.

[0164] S503: The server performs personalized learning on the user preference of the terminal device to obtain second preference information, and the terminal device performs personalized learning on the user preference of the terminal device to obtain third preference information.

[0165] The terminal device may send its user data to the server. The server then learns the user data based on the first model to obtain second preference information, which is the coarse-grained features learned by the first model. The terminal device may also learn the user data based on the second model to obtain third preference information, which is the fine-grained features learned by the second model. In some embodiments, the server may return the second preference information to the terminal device.

[0166] In some embodiments, the server may input the user data of the terminal device into the first model to obtain the second preference information output by the first model. In some embodiments, the second preference information may indicate the AI ​​model preferred by the user based on the business type corresponding to the AI ​​model and the application to which the AI ​​model belongs, or in other words, each AI model corresponds to a business type and an application. Exemplarily, the super-resolution model belonging to video application A and the super-resolution model belonging to video application B can be used as two AI models, thereby indicating the user's preference for the super-resolution model belonging to video application A and the user's preference for the super-resolution model belonging to video application B, respectively.

[0167] In some embodiments, the second preference information may include a second vector, the second vector including at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

[0168] For example, the second vector may be [0.1, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.2]. Each element may correspond to an AI model, and the value of each element is greater than or equal to 0 and less than or equal to 1. The closer the value of each element is to 1, the more the user prefers to use the AI ​​model corresponding to the element.

[0169] Alternatively, in other embodiments, the second preference information may indicate the user's preferred AI model based on the business type corresponding to the AI ​​model. For example, the super-resolution model belonging to video application A and the super-resolution model belonging to video application B may be considered as one AI model, thereby indicating the user's preference for the super-resolution model.

[0170] In some embodiments, the server may input the user data of the terminal device into the first model to obtain the fourth preference information output by the first model, and the server compresses the fourth preference information to obtain the second preference information. Among them, the fourth preference information may indicate the user's preferred AI model based on the business type corresponding to the AI ​​model, the application to which the AI ​​model belongs, and more dimensions for distinguishing AI models, thereby more accurately indicating the user's preference for the AI ​​model based on multiple dimensions. In some embodiments, the fourth preference information may indicate the user's preferred AI model based on the device type where the AI ​​model is located, the business type corresponding to the AI ​​model, and the application to which the AI ​​model belongs, or in other words, each AI model corresponds to a terminal device, a business type, and an application. Exemplarily, the ASR model deployed in the car computer and the ASR model deployed in the mobile phone can be used as two AI models, thereby respectively indicating the user's preference for the ASR model deployed in the car computer and the user's preference for the ASR model deployed in the mobile phone.

[0171] In some embodiments, the fourth preference information may include a fourth vector, the fourth vector including at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

[0172] In some embodiments, the terminal device may input the user data of the terminal device into the second model to obtain third preference information output by the second model. In some embodiments, the third preference information may include a third vector, the third vector including at least one element, each element of the at least one element corresponding to an AI model, and the value of each element indicating the user's preference for the AI ​​model corresponding to the element. In some embodiments, the third preference information may indicate the user's preferred AI model based on the business type corresponding to the AI ​​model and the application to which the AI ​​model belongs.

[0173] Alternatively, in other embodiments, the third preference information may indicate the user's preferred AI model based on the business type corresponding to the AI ​​model.

[0174] S504: The terminal device merges the second preference information and the third preference information to obtain the first preference information.

[0175] By fusing the second preference information learned by the server and the third preference information learned by the terminal device to generate the first preference information, the long-tail user problem of the model when the server learns the user's preference for the AI ​​model can be improved, as well as the overfitting problem caused by the terminal device due to less user data, thereby improving the accuracy of learning the user's preference for the AI ​​model.

[0176] In some embodiments, the second preference information includes a second vector, and the third preference information includes a third vector. The terminal device may add the product of the second vector and the first coefficient and the product of the third vector and the second coefficient to obtain a first vector. The first vector is the first preference information, and the first vector includes at least one element, each of which corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

[0177] The sum of the first coefficient and the second coefficient is 1. It is understood that the embodiments of the present application do not limit the method for setting the first coefficient and the second coefficient, nor the specific values ​​of the first coefficient and the second coefficient. For example, the first coefficient and the second coefficient can be set in advance by relevant technical personnel, and the first coefficient and the second coefficient are both 0.5.

[0178] For example, the second vector is [0.1, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.2], the third vector is [0.2, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.3], and the first coefficient and the second coefficient are both 0.5. Then the first vector = second vector * 0.5 + third vector * 0.5, that is, the first vector is [0.15, 0.2, 0.5, 0.1, 0.11, 0.6, 0.9, 0.25].

[0179] It is understood that the method for fusing the second preference information and the third preference information may correspond to the data types of the second preference information and the third preference information. When the second preference information and the third preference information are information of other data types, the terminal device may also fuse the second preference information and the third preference information using a fusing method corresponding to the other data type. For example, if the data type of the second preference information and the third preference information is key-key-value, the terminal device may add the key values ​​corresponding to the same key to obtain a new key value corresponding to the key.

[0180] In an embodiment of the present application, both the terminal device and the server can learn the terminal device's preference for the AI ​​model based on the user data of the terminal device, and obtain second preference information and third preference information respectively, and then fuse the second preference information and the third preference information to obtain the first preference information, which improves the long-tail user problem of the model when the server learns the user's preference for the AI ​​model, and improves the overfitting problem of the terminal device due to less user data, thereby improving the accuracy of learning the user's preference for the AI ​​model.

[0181] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application provides a device for deploying an AI model. The device embodiment corresponds to the above method embodiment. For ease of reading, this device embodiment will no longer repeat the details of the above method embodiment one by one, but it should be clear that the device in this embodiment can correspond to and implement all the contents of the above method embodiment.

[0182] In one possible implementation, the AI ​​model deployment device may correspond to the terminal device in the above method embodiment, for example, the terminal device or a chip configured in the terminal device. The AI ​​model deployment device is used to execute each step or process corresponding to the terminal device in the above method.

[0183] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application provides a device for deploying an AI model. The device embodiment corresponds to the above method embodiment. For ease of reading, this device embodiment will no longer repeat the details of the above method embodiment one by one, but it should be clear that the device in this embodiment can correspond to and implement all the contents of the above method embodiment.

[0184] In one possible implementation, the AI ​​model deployment device may correspond to the terminal device in the above method embodiment, for example, a server or a chip configured in a server. The AI ​​model deployment device is used to execute the various steps or processes corresponding to the server in the above method.

[0185] Based on the same inventive concept, an embodiment of the present application further provides a terminal device comprising: a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method described in the above method embodiment when calling the computer program.

[0186] The terminal device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be repeated here.

[0187] Based on the same inventive concept, an embodiment of the present application further provides a chip system, which includes a processor coupled to a memory, and executes a computer program stored in the memory to implement the method described in the above method embodiment.

[0188] The chip system may be a single chip or a chip module composed of multiple chips.

[0189] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the above method embodiment is implemented.

[0190] An embodiment of the present application further provides a computer program product, which, when executed on a terminal device, enables the terminal device to implement the method described in the above method embodiment.

[0191] Based on the same inventive concept, an embodiment of the present application further provides a server comprising: a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method described in the above method embodiment when the computer program is called.

[0192] The server provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0193] Based on the same inventive concept, an embodiment of the present application further provides a chip system, which includes a processor coupled to a memory, and executes a computer program stored in the memory to implement the method described in the above method embodiment.

[0194] The chip system may be a single chip or a chip module composed of multiple chips.

[0195] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the above method embodiment is implemented.

[0196] An embodiment of the present application further provides a computer program product, which, when executed on a server, enables the server to implement the method described in the above method embodiment.

[0197] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include at least: any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.

[0198] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0199] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0200] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0201] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0202] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0203] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0204] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0205] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for deploying an artificial intelligence model, characterized in that: The method comprises: The terminal device sends hardware information, resource information and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; The server acquires, based on the hardware information, the resource information, and the first preference information, a first operator library matching the terminal device, where the first operator library includes at least one operator; The server returns the first operator library to the terminal device; The terminal device deploys the AI ​​model based on the first operator library.

2. The method according to claim 1, characterized in that The method further comprises: The terminal device sends user data of the terminal device to the server; The server obtains second preference information based on the user data; The server returns the second preference information to the terminal device; The terminal device acquires third preference information based on the user data; The terminal device generates the first preference information based on the second preference information and the third preference information.

3. The method according to claim 1, characterized in that The method further comprises: The terminal device acquires third preference information based on user data of the terminal device; The terminal device sends the user data and the third preference information to the server; The server obtains second preference information based on the user data; The server generates the first preference information based on the second preference information and the third preference information.

4. The method according to claim 2 or 3, characterized in that: The server obtains second preference information based on the user data, including: The server inputs the user data into the first model to obtain the second preference information output by the first model; The terminal device acquires third preference information based on the user data, including: The terminal device inputs the user data into a second model to obtain the third preference information output by the second model, where the second model is obtained by compressing the first model.

5. The method according to claim 4, characterized in that The method further comprises: The server obtains the first model based on at least two user data; The server compresses the first model to obtain the second model; The server returns the second model to the terminal device.

6. The method according to any one of claims 2 to 5, characterized in that: The first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

7. The method according to claim 6, characterized in that The second preference information includes a second vector, the third preference information includes a third vector, and the terminal device generates the first preference information based on the second preference information and the third preference information, including: The terminal device adds the product of the second vector and the first coefficient and the product of the third vector and the second coefficient to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.

8. The method according to any one of claims 1 to 7, characterized in that: The terminal device deploys the AI ​​model based on the first operator library, including: The terminal device compiles the AI ​​model in the intermediate representation IR format based on the first operator library to obtain the AI ​​model in the executable file format; and / or, The terminal device runs the AI ​​model in the executable file format.

9. The method according to any one of claims 1 to 7, characterized in that: The terminal device further includes a preset second operator library, the second operator library includes at least one operator, and the terminal device deploys the AI ​​model based on the first operator library, including: The terminal device deploys the AI ​​model based on the first operator library and the second operator library.

10. A method for deploying an artificial intelligence model, characterized in that: Applied to a terminal device, the method comprises: Sending hardware information, resource information, and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; Acquire a first operator library returned by the server, where the first operator library includes at least one operator; The terminal device deploys the AI ​​model based on the first operator library.

11. The method according to claim 10, characterized in that The method further comprises: Sending user data of the terminal device to the server; Acquire second preference information returned by the server based on the user data; acquiring third preference information based on the user data; The first preference information is generated based on the second preference information and the third preference information.

12. The method according to claim 11, characterized in that The first preference information includes a first vector, the first vector includes at least one element, each element of the at least one element corresponds to an AI model, and the value of each element indicates the user's preference for the AI ​​model corresponding to the element.

13. The method according to claim 12, characterized in that The second preference information includes a second vector, the third preference information includes a third vector, and generating the first preference information based on the second preference information and the third preference information includes: The product of the second vector and the first coefficient and the product of the third vector and the second coefficient are added to obtain the first vector, and the sum of the first coefficient and the second coefficient is 1.

14. A method for deploying an artificial intelligence model, characterized in that: Applied to a server, the method comprises: Acquire hardware information, resource information, and first preference information of a terminal device, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by a user of the terminal device; Based on the hardware information, the resource information, and the first preference information, acquiring a first operator library matching the terminal device, the first operator library including at least one operator; The first operator library is returned to the terminal device, where the first operator library is used for the terminal device to deploy the AI ​​model.

15. The method according to claim 14, characterized in that The method further comprises: Acquiring user data and third preference information of the terminal device; acquiring second preference information based on the user data; The first preference information is generated based on the second preference information and the third preference information.

16. The method according to claim 15, characterized in that The method further comprises: acquiring a first model based on at least two user data, wherein the first model is used by the server to acquire the third preference information; compressing the first model to obtain a second model, where the second model is used by the terminal device to obtain the second preference information; The second model is returned to the terminal device.

17. A system, characterized in that: The system includes a server and a terminal device; The terminal device is used to send hardware information, resource information and first preference information of the terminal device to the server, wherein the hardware information is used to indicate the hardware performance of the terminal device, the resource information is used to indicate the status of the device resources of the terminal device, and the first preference information is used to indicate at least one artificial intelligence AI model preferred by the user of the terminal device; deploying the AI ​​model based on a first operator library, wherein the first operator library includes at least one operator; The server is used to obtain a first operator library matching the terminal device based on the hardware information, the resource information and the first preference information; and return the first operator library to the terminal device.

18. A terminal device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method according to any one of claims 10 to 13 when calling the computer program.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 10 to 13 is implemented.

20. A computer program product, characterized in that When the computer program product is executed on a terminal device, the terminal device is enabled to execute the method according to any one of claims 10 to 13.

21. A server, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the method according to any one of claims 14 to 16 when calling the computer program.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 14 to 16 is implemented.

23. A computer program product, characterized in that When the computer program product is run on a server, the server is caused to execute the method according to any one of claims 14 to 16.

Citation Information

Patent Citations

  • Model using method and device

    CN111753999A

  • Method and device for generating information

    CN111784377A

  • Neural network computing deployment method and device, storage medium and computer equipment

    CN112686378A

  • Downstream operator recommendation method, electronic equipment and computer readable storage medium

    CN115167846A

  • Weak supervised learning driven large and small model coevolution method and terminal

    CN116994096A

Cited By

  • Operator library generation method for coarseness reconfigurable AI array and application

    CN120525013A

  • Heterogeneous resource type selection method and device

    CN121501857A