Task execution methods, plug-ins, devices, media, and products
By embedding alternative task execution units in terminal devices and performing format conversion and task scheduling, the problem of inference task execution caused by the fragmentation of the terminal device's operating environment is solved, and intelligent functions are compatible and executed normally on multiple terminal devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 阿里巴巴(中国)网络技术有限公司
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-26
AI Technical Summary
How to ensure the normal execution of inference tasks in different types and operating environments of terminal devices, especially to realize intelligent functions in fragmented terminal device operating environments.
A task execution plugin is provided, which has built-in alternative task execution units. It determines the target task execution unit based on the operating system and system-on-a-chip type of the terminal device, and ensures that the inference task is executed in an appropriate operating environment through format conversion and task scheduling.
It achieves compatibility of inference tasks in various terminal device operating environments, reduces the difficulty of implementing intelligent functions on terminal devices, and ensures the normal execution of intelligent functions in fragmented environments.
Smart Images

Figure CN122086489A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a task execution method, plug-in, device, medium, and product. Background Technology
[0002] With the development of artificial intelligence technology, intelligent functions such as speech recognition, intelligent question answering, personalized recommendations, and real-time translation can be introduced into terminal devices. These intelligent functions can typically be executed within the terminal device as inference tasks to ensure their proper functioning. These inference tasks can be executed within the runtime environment provided by the terminal device.
[0003] However, in reality, there are more and more types of terminal devices, such as mobile phones, tablets, smart wearable devices, etc., and the operating environments provided by different types of terminal devices are also different. At this time, how to ensure that inference tasks can be executed normally in different operating environments has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of this application provide a task execution method, plug-in, device, medium, and product to ensure that inference tasks can be executed normally in different operating environments.
[0005] This application provides a task execution method, the method comprising: Acquire inference tasks generated using the intelligent functions on the terminal device; Among the alternative task execution units provided by the task execution plugin built into the terminal device, the target task execution unit is determined. The different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device. The inference task is executed using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task.
[0006] This application provides a task execution plugin, which is configured to perform the following operations: Acquire inference tasks generated using the intelligent functions on the terminal device; Among the alternative task execution units provided by the task execution plugin built into the terminal device, the target task execution unit is determined. The different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device. The inference task is executed using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task.
[0007] This application provides a task execution method, the method comprising: Acquire inference tasks generated using the intelligent functions on the terminal device; Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is format-converted to obtain the target inference model; The inference task is performed using the target inference model running on a neural network processor provided by the terminal device.
[0008] This application provides a task execution plugin, which is configured to perform the following operations: Acquire inference tasks generated using the intelligent functions on the terminal device; Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is format-converted to obtain the target inference model; The inference task is performed using the target inference model running on a neural network processor provided by the terminal device.
[0009] This application provides an electronic device including a processor and a memory. The memory stores one or more computer instructions, which, when executed by the processor, implement the task execution method described above. The electronic device may also include a communication interface for communicating with other devices or communication networks.
[0010] This application provides a non-transitory machine-readable storage medium storing executable code. When the executable code is executed by a processor of an electronic device, the processor can at least implement the task execution method described above.
[0011] This application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the task execution method described above.
[0012] The task execution method provided in this application includes: after a user uses the intelligent functions on a terminal, the terminal device can generate an inference task. Then, a task execution plugin installed on the terminal device can determine a target task execution unit from among its provided candidate task execution units to execute the inference task. Different task execution units among the candidate units are suitable for different operating environments provided by different terminal devices, and different task execution units correspond to different intelligent functions. Finally, the inference task generated by the terminal device can be executed in the target operating environment using the target task execution unit in the task execution plugin.
[0013] Because the task execution plugin has built-in task execution units suitable for different operating environments—that is, alternative task execution units—it can be considered a plugin compatible with multiple operating environments. Due to this compatibility, once the plugin is installed on a terminal device and an inference task is generated, regardless of the terminal device's operating environment, the task execution plugin can determine the appropriate task execution unit from the alternative units, and that unit will ensure the normal execution of the inference task.
[0014] In addition, the compatibility of the task execution plugin can ensure that the inference task can be executed normally in fragmented operating environments (i.e., multiple operating environments). At the same time, it can also reduce the difficulty of implementing intelligent functions on terminal devices. That is, there is no need to develop corresponding modules for executing inference tasks for different operating environments, and the terminal device can provide rich intelligent functions. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of a task execution method provided in an embodiment of this application; Figure 2 A flowchart illustrating another task execution method provided in this application embodiment; Figure 3 A flowchart illustrating the process of determining a target task execution unit, as provided in an embodiment of this application; Figure 4 A flowchart illustrating yet another task execution method provided in this application embodiment; Figure 5 A schematic diagram illustrating the process of obtaining the second target model provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a task execution plugin provided in an embodiment of this application; Figure 7This is a schematic diagram of the architecture of a terminal device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of another task execution plugin provided in an embodiment of this application; Figure 9 A schematic diagram of the architecture of a task execution plugin provided in an embodiment of this application; Figure 10 A flowchart illustrating yet another task execution method provided in this application embodiment; Figure 11 This is a schematic diagram of the structure of another task execution plugin provided in an embodiment of this application; Figure 12 This is an application diagram of a cloud computing environment provided in an embodiment of this application; Figure 13 This application provides a schematic diagram of the structure of an electronic device. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “said,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.
[0018] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0019] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to identification.” Similarly, depending on the context, the phrases “if determination” or “if identification (of the condition or event of the statement)” can be interpreted as “when determination” or “in response to determination” or “when identification (of the condition or event of the statement)” or “in response to identification (of the condition or event of the statement).”
[0020] It should be noted that, in the cases involving user interaction operations or triggering operations in the embodiments of this application, the user interaction operations or triggering operations involved in the embodiments of this application include, but are not limited to, various interaction operations such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations; among them, touch operations include, but are not limited to, click operations, double click operations, long press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight line swipes and curved line swipes.
[0021] It should be noted that, in the case of user information involved in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0022] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.
[0023] In addition, before the detailed description of each of the following embodiments of this application, the background information can be used to introduce the application's usage background: With the continuous development of artificial intelligence technology, intelligent functions such as image recognition, voice interaction, intelligent question answering, image search, text search, and personalized recommendations can be integrated into terminal devices, thereby improving the intelligence level of these devices. These terminal devices can be various types, including mobile phones, tablets, smart TVs, and smart wearable devices.
[0024] Optionally, smart wearable devices may specifically include smartwatches, artificial intelligence (AI) glasses, augmented reality (AR) display devices, virtual reality (VR) display devices, and so on. Optionally, the intelligent functions can be provided by the terminal device itself; for example, AI glasses can provide intelligent functions such as image recognition and voice interaction; VR display devices can provide intelligent functions such as voice interaction and personalized recommendations. Optionally, intelligent functions can also be provided by applications (APPs) installed on the terminal device, such as image search, text search, and personalized recommendations provided in e-commerce apps.
[0025] Furthermore, the aforementioned intelligent functions can be implemented using rule-based algorithms, and more commonly, they can be implemented using inference models. When intelligent functions are implemented using inference models, specifically, the inference model can execute inference tasks generated after the user triggers the intelligent function within the operating environment provided by the terminal device, thereby realizing the intelligent function. However, in reality, not only are there diverse types of terminal devices, but the operating environments they provide are also varied. In other words, both the types of terminal devices and their operating environments are fragmented. Therefore, it is necessary to ensure that the inference model can run normally in the fragmented operating environments of the terminal devices (i.e., the different operating environments provided by different types of terminal devices) to execute inference tasks. In this case, the methods and plugins provided in the following embodiments of this application can be used.
[0026] The operating environment provided by the terminal device may include the terminal system's operating system and / or the terminal device's system-on-chip (SoC). Optionally, the operating system of a mobile phone may include Android, iOS, HarmonyOS, etc., while the operating system of a smartwatch may include WatchOS, LiteOS, etc. Optionally, the SoC may include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Neural Processing Unit (NPU), a Tensor Processing Unit (TPU), etc. It is easy to understand that a terminal device can have one operating system, a terminal device can have at least one SoC, and different SoCs can handle different inference tasks.
[0027] It should be noted that the inference models mentioned in the following embodiments of this application can be deep learning models with relatively large parameter sizes. Furthermore, the embodiments of this application do not limit the number of model parameters supported by the models mentioned in each embodiment, aiming to meet actual needs. If there are relatively many model parameters, the model size will be relatively large, and the model performance will be relatively better, but it will consume more time and resources during inference or training. Conversely, if there are relatively few model parameters, the model size will be relatively small, and while meeting performance requirements, the model will be more lightweight, consuming less time and resources during inference or training. The inference models involved in the embodiments of this application may optionally be language models (LM) used to process and generate natural language text, or multimodal models (MM) used to process multimodal data, etc. Essentially, they are deep learning models, which can be implemented based on neural network architectures and can be obtained through pre-training with large amounts of data. In one alternative implementation, the inference model may include an encoder, a decoder, a self-attention layer, and a feed-forward neural network. The encoder converts the input data (usually in sequence) into a vector representation, capturing the semantic features of the input data. The decoder transforms the intermediate representation generated by the encoder into output data (usually in sequence). The self-attention layer is a mechanism that allows the model to focus on other positions in the sequence to better encode information about the current position. The feed-forward neural network can perform nonlinear transformations on the output of the self-attention layer to enhance the model's expressive power. These components work together to enable models built upon them to perform well in various complex inference tasks, such as natural language processing, image recognition, speech recognition, intelligent question answering, image search, and text search.
[0028] Based on the above, some embodiments of this application will be described in detail below with reference to the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0029] Figure 1 This is a flowchart illustrating a task execution method provided in an embodiment of this application. This method can be executed by a task execution plugin installed on a terminal device. Figure 1 As shown, the method may include the following steps: S101, Obtain the inference task generated using the intelligent functions on the terminal device.
[0030] As described in the background section above, the terminal device itself or the app installed on it can provide various intelligent functions. When a user triggers a function on the terminal device, the device can generate a corresponding inference task, which can then be obtained by a task execution plugin installed on the device. Optionally, this task execution plugin can be integrated into the app installed on the terminal device, or it can be installed separately by the user as a third-party plugin.
[0031] For example, when the terminal device is AI glasses in a smart wearable device, the user can generate voice commands to the AI glasses. This is equivalent to activating the AI glasses' voice recognition function, and the inference task generated by the AI glasses can specifically be a voice recognition task. The voice commands generated by the user can be used to control the AI glasses to play music, translate text, and so on.
[0032] When the terminal device is a mobile phone or tablet with an e-commerce app installed, the user can initiate the app launch, which can be considered a usage operation triggering the personalized recommendation function. Therefore, the inference task generated by the terminal device can specifically be a personalized recommendation task. Subsequently, the user can also enter images as search keywords on the pages provided by the e-commerce app and click the search component. These input and search operations are equivalent to triggering the image search function, so the inference task generated by the terminal device can specifically be an image search task. Similarly, the inference tasks generated when the user uses the e-commerce app can also include intelligent question-answering tasks, which can be generated during the pre-sales and after-sales stages of the product.
[0033] S102, determine the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device. Different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device.
[0034] S103, in the target operating environment provided by the terminal device that generates the inference task, the inference task is executed using the target task execution unit.
[0035] After acquiring the inference task, the task execution plugin can utilize its capability detection function to determine the target task execution unit from among its provided candidate task execution units. Finally, within the target runtime environment provided by the terminal device, the inference task is executed using this target task execution unit. The capability detection function of the task execution plugin is essentially a task scheduling function, which can be used to schedule the inference task to the target task execution unit. Optionally, this task scheduling function can be implemented by the task scheduling unit within the task execution plugin.
[0036] Among the candidate task execution units, each task execution unit may be applicable to a different operating environment. Optionally, any task execution unit among the candidate task execution units may be applicable to at least one operating environment. The operating environment may include the operating system and / or system-on-a-chip of the terminal device, and the specific contents of the operating system and system-on-a-chip can be found in the relevant introduction in the above-mentioned background of use, and will not be repeated here. Optionally, the number of target task execution units determined by the task execution plugin from the candidate task execution units for executing the inference task generated using any intelligent function may be at least one.
[0037] In this embodiment, after the user uses the intelligent functions on the terminal, the terminal device can generate an inference task. Then, the task execution plugin installed on the terminal device can determine the target task execution unit from among its provided candidate task execution units to perform the inference task. Different task execution units among the candidate units are suitable for different operating environments. Finally, the target task execution unit in the task execution plugin can be used to execute the inference task generated by the terminal device in the target operating environment.
[0038] Because the task execution plugin has built-in task execution units suitable for different operating environments—that is, alternative task execution units—it can be considered a plugin compatible with multiple operating environments. Due to this compatibility, once the plugin is installed on a terminal device and an inference task is generated, regardless of the terminal device's operating environment, the task execution plugin can determine the appropriate task execution unit from the alternative units, and that unit will ensure the normal execution of the inference task.
[0039] In addition, the compatibility of the task execution plugin can ensure that the inference task can be executed normally in fragmented operating environments (i.e., multiple operating environments). At the same time, it can also reduce the difficulty of implementing intelligent functions on terminal devices. That is, there is no need to develop corresponding modules for executing inference tasks for different operating environments, and the terminal device can provide rich intelligent functions.
[0040] Figure 2A flowchart illustrating another task execution method provided in this application embodiment. This method can also be executed by a task execution plugin installed on a terminal device. Figure 2 As shown, the method may include the following steps: S201, Obtain the inference task generated using the intelligent functions on the terminal device.
[0041] The specific implementation process of step S201 in this embodiment can be found in [reference needed]. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0042] S202, determine the first acquisition interface provided by the task execution plugin based on the type of inference task.
[0043] S203, call the first acquisition interface to obtain the task initialization pipeline corresponding to the inference task.
[0044] S204, Run the task initialization pipeline to determine the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device. Different task execution units among the alternative task execution units are suitable for different operating environments, including the terminal device's operating system and / or system-on-a-chip.
[0045] S205, Run the task initialization pipeline to create the execution environment for the inference task in the target runtime environment.
[0046] S206, In the execution environment, the target task execution unit is used to perform the inference task.
[0047] Subsequently, the task execution plugin can determine the first acquisition interface corresponding to the inference task type based on the task type. Calling this first acquisition interface allows the plugin to obtain the task initialization pipeline corresponding to the inference task. This pipeline describes various initialization processes required before the inference task execution, such as determining the target task execution unit from candidate task execution units and creating the execution environment for the target task execution unit within the target runtime environment provided by the terminal device. Finally, after obtaining the task initialization pipeline through the interface call, the task execution plugin can further execute steps S204-S205, that is, run the task initialization pipeline to obtain the target task execution unit and the execution environment for the inference task. Ultimately, the task execution plugin can then execute the inference task within the execution environment using the target task execution unit.
[0048] Optionally, the type of reasoning task can also be considered as the content of the reasoning task, just as... Figure 1The examples shown in the embodiments include inference tasks generated when using e-commerce apps, such as speech recognition tasks, personalized recommendation tasks, image search tasks, and intelligent question answering tasks, which are all different types of tasks.
[0049] Optionally, the task execution plugin can be configured with multiple acquisition interfaces, each of which can correspond to at least one type of inference task. Therefore, by calling different acquisition interfaces, the task initialization pipelines corresponding to different types of inference tasks can be obtained. For example, the ASRPipeline corresponding to the Automatic Speech Recognition (ASR) task can be obtained, the Large Language ModelPipeline corresponding to the intelligent question answering task can be obtained, and the ObjectDetect Pipeline corresponding to the image search task can be obtained, etc.
[0050] In this embodiment, after obtaining the inference task, the task execution plugin can obtain the task initialization pipeline corresponding to the inference task through an interface call, and prepare for the execution of the inference task by running the pipeline. Then, the inference task can be executed in the execution environment using the target task execution unit, thereby achieving the normal execution of the inference task. Furthermore, for details not described in this embodiment and the technical effects achieved, please refer to [link to relevant documentation]. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0051] As described in the above embodiments, to ensure the normal execution of the inference task in a fragmented runtime environment, the task execution plugin involves a process of "using its own task scheduling function to query the target task execution unit from the candidate task execution units and scheduling the inference task to that target task execution unit." The query process of the task execution plugin, that is, the determination process of the target task execution unit, will be described in detail below.
[0052] Optionally, the reasoning tasks that different task execution units in the candidate task units can each execute can correspond to different smart devices. Continuing with the example in step S101, the speech recognition task, personalized recommendation task, image search task, and intelligent question answering task can be executed by different task execution units in the candidate task units. Furthermore, the different task execution units in the candidate task units can have different presentation formats, different deployment locations, and different query priorities.
[0053] Optionally, the alternative task execution units in the task execution plugin may include alternative inference interfaces. Different inference interfaces correspond to different operating systems of the terminal device. Therefore, different inference interfaces can be used to call execution components in terminal devices with different operating systems, allowing those execution components to specifically execute the inference task. The execution component may include a first inference model and / or a rule algorithm, and the alternative inference interfaces have the highest query priority. Furthermore, the alternative inference interfaces can be integrated into the first execution module of the task execution plugin.
[0054] Optionally, the alternative task execution unit may further include a second inference model, which can be deployed locally on the terminal device. Optionally, this second inference model can be invoked by the second execution module in the task execution plugin, and the invoked inference model is used to execute the inference task. The query priority of the second inference model is lower than that of the first inference model, and the number of parameters in the second inference model is greater than that in the first inference model. Therefore, compared to using the first inference model to execute the inference task, using the second inference model involves a larger computational load, especially for inference tasks generated by the APP. If the second inference model is used to execute the inference task, the larger computational load may also affect the response speed of other functions of the APP.
[0055] In addition, the inference tasks performed by the execution components called by the alternative inference interface and the inference tasks performed by the second inference model can correspond to different intelligent functions. Therefore, the alternative inference interface and the second inference model can complement each other, thereby enriching the intelligent functions that the terminal device can realize.
[0056] For the two query priorities mentioned above, the query process of the task execution plugin can be described as follows: Once the task execution plugin obtains the inference task, it can first determine whether there is a target inference interface among the candidate inference interfaces based on the type of inference task. This target inference interface is obviously corresponding to the intelligent function.
[0057] In one scenario, if the target inference interface exists among the candidate inference interfaces, the task execution plugin can directly execute the inference task in the target runtime environment using the execution component called by the target inference interface. In another scenario, if the target inference interface does not exist among the candidate inference interfaces, the task execution plugin can determine whether a first target model corresponding to the intelligent function exists in the second inference model based on the type of the inference task.
[0058] In one scenario, if the first target model exists within the second inference model, the task execution plugin can directly execute the inference task using the first target model within the target runtime environment. In another scenario, if the first target model does not exist within the second inference model—meaning all candidate task execution units in the task execution plugin have been queried but no target task execution unit has been found—it indicates that the terminal device cannot execute the inference task. In this case, the task execution plugin can output a notification message reflecting the task execution failure, which can also be displayed to the user via the terminal device.
[0059] In this embodiment, on the one hand, for the candidate inference interfaces and the second inference model with decreasing query priority, after generating the inference task, the task execution plugin can first query the candidate inference interfaces. If the target task execution unit is not found, it will then query the target task execution unit in the second execution module. This ensures the normal execution of the inference task while improving query speed and reducing query costs. Furthermore, the two-level task execution unit setup can enrich the intelligent functions provided by the terminal device.
[0060] On the other hand, since the different alternative inference interfaces in the alternative task execution units correspond to different operating systems of the terminal devices (that is, to fragmented operating systems), the compatibility of the task execution plugin can be specifically reflected in its compatibility with the terminal device operating system. This compatibility can ensure that the inference tasks can be executed normally in fragmented operating systems, which improves the problem of difficulty in implementing intelligent functions on the terminal device due to the fragmentation of the terminal device operating system.
[0061] Based on the aforementioned two-level alternative task execution units, optionally, the alternative task execution units may further include a third inference model deployed in the task execution plugin. The third inference model can be called by the third execution module in the task execution plugin. The number of parameters of this third inference model is between that of the first and second inference models. The intelligent functions that the third inference model can implement can also differ from those implemented by the second inference model and the alternative inference interface; therefore, the configuration of the third inference model can also enrich the intelligent functions provided by the terminal device. Optionally, the query priority of this third inference model is lower than that of the first inference model but higher than that of the second inference model.
[0062] For the three query priorities of the candidate task execution units mentioned above, the query process of the task execution plugin can be described as follows: Once the task execution plugin obtains the inference task, it can use the first execution module, which means it can first determine whether there is a target inference interface corresponding to the intelligent function among the candidate inference interfaces, based on the type of inference task.
[0063] In one scenario, if the target inference interface exists among the candidate inference interfaces, the task execution plugin can directly execute the inference task within the target runtime environment using the execution component called by the target inference interface. In another scenario, if the target inference interface does not exist among the candidate inference interfaces, the task execution plugin can determine, based on the third inference model, whether a second target model corresponding to the intelligent function exists.
[0064] In one scenario, if a second target model exists, the task execution plugin can directly execute the inference task within the target runtime environment using the second target model. In another scenario, if a second target model does not exist, the task execution plugin can determine, within the second inference model, whether a first target model corresponding to the intelligent function exists.
[0065] In one scenario, if the first target model exists within the second inference model, the task execution plugin can directly execute the inference task using the first target model within the target runtime environment. In another scenario, if the first target model does not exist within the second inference model, it indicates that the terminal device cannot execute the inference task. In this case, the task execution plugin can output a notification message indicating task execution failure, which can also be displayed to the user using the terminal device's display function.
[0066] In this embodiment, the query priorities of the alternative inference interface, the third inference model, and the second inference model decrease sequentially. Therefore, after obtaining the inference task, the task execution plugin can query sequentially according to the query priority to ensure the normal execution of the inference task using the queried target task execution unit. Furthermore, similar to the above embodiment, the multi-level execution unit setup can enrich the intelligent functions provided by the terminal device, improve search speed, reduce search costs, and ensure the compatibility of the task execution plugin, thus guaranteeing the normal execution of the inference task even in fragmented operating systems.
[0067] Based on the aforementioned three-level query priority candidate execution units, the candidate task execution units may optionally include a fourth inference model deployed in a cloud server. This fourth inference model can be called by the fourth execution module of the task execution plugin, and the fourth inference model has the lowest query priority.
[0068] For the four query priorities of the candidate task execution units mentioned above, the query process of the task execution plugin can be described as follows: Once the reasoning task is obtained, if the task execution plugin fails to find a target task execution unit for executing the reasoning task during the three rounds of queries mentioned in the above embodiments, the task execution plugin can determine whether there is a third target model corresponding to the intelligent function in the fourth reasoning model.
[0069] In one scenario, if the third target model exists within the fourth inference model, the task execution plugin can directly execute the inference task within the target runtime environment using the third target model. In another scenario, if the third target model does not exist within the fourth inference model, it indicates that the terminal device cannot execute the inference task. In this case, the task execution plugin can output a notification message indicating task execution failure, which can also be displayed to the user using the terminal device's display function.
[0070] The query process for the fourth-level candidate task execution unit in this embodiment can also be found in [reference needed]. Figure 3 The flowchart shown is, in other words, the process is Figure 1 The specific execution process of step S102 in the illustrated embodiment.
[0071] In this embodiment, the query priorities of the alternative inference interface, the third inference model, the second inference model, and the fourth inference model decrease sequentially. Therefore, after obtaining the inference task, the task execution plugin can query sequentially according to the query priority to ensure the normal execution of the inference task using the queried target task execution unit. Furthermore, similar to the above embodiment, the multi-level execution unit setup can enrich the intelligent functions provided by the terminal device, improve search speed, reduce search costs, and ensure the compatibility of the task execution plugin, thus guaranteeing the normal execution of the inference task even in fragmented operating systems.
[0072] In summary, the alternative task execution units in the task execution plugin can have four query priorities, and the intelligent functions corresponding to the inference tasks executed by different task execution units can be different. Therefore, the setting of multi-level alternative task execution units can enrich the intelligent functions provided by the terminal device. At the same time, the hierarchical query process of the task execution plugin for alternative execution units can improve search speed and reduce search costs.
[0073] Optionally, the query priority of the alternative task execution units can be set based on the commonness of the applicable operating environment. For example, the target inference interface with the highest query priority can run the first inference model and / or rule algorithm obtained by calling this interface in a relatively common operating environment, such as a common system-on-a-chip (SoC) like a CPU or GPU, and the terminal system's operating system can also be any common system such as Android, iOS, or HarmonyOS. The third inference model with the second highest query priority can run on a common SoC like a CPU or GPU, or on a less common SoC like an NPU. The fourth inference model with the lowest query priority can run on an uncommon operating system.
[0074] Furthermore, since the alternative inference interface is directly deployed on the terminal device side, the third and second inference models can be directly invoked on the terminal device side. Therefore, if the target task execution unit is obtained based on the alternative inference interface, the third inference model, and the second inference model, the execution of the inference task can be performed directly on the terminal side without using the network. In this case, even if the terminal device is offline, the normal provision of intelligent functions can be guaranteed.
[0075] Since the fourth inference model can be deployed on cloud servers, if the target task execution unit is obtained from the fourth inference model, the inference task needs to be executed in the cloud via the network. However, because the fourth inference model is applicable to uncommon operating systems, even terminal devices with uncommon operating systems can provide intelligent functions normally using the fourth inference model.
[0076] In the above embodiments, the compatibility of the task execution plugin can ensure that the inference task can be executed normally in a fragmented runtime environment, especially on a fragmented operating system. However, in practice, since the system-on-a-chip (SoC) provided by the terminal device can be diverse, that is, the SoC of the terminal device can also be fragmented, this application can also be used. Figure 4 The embodiments shown are designed to ensure that inference tasks can be executed normally on fragmented system-on-a-chip.
[0077] Figure 4 This is a flowchart illustrating yet another task execution method provided in an embodiment of this application. Figure 4 The method described above can be used under the following conditions: the system-on-a-chip provided by the terminal device generating the inference task can be a faster NPU or TPU, rather than a common CPU or GPU, and the task execution plugin does not find the target task unit in the alternative inference interface, but instead finds the target task execution unit based on the third inference model. In this case, such as Figure 4 As shown, the method may include the following steps: S301, Obtain the inference task generated using the intelligent functions on the terminal device.
[0078] The specific implementation process of step S301 in this embodiment can be found in [reference needed]. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated in this step.
[0079] S302, Based on the third reasoning model, determine whether there exists a second target model corresponding to the intelligent function.
[0080] Alternatively, the system-on-a-chip (SoC) used in the third inference model may be the same as the SoC provided by the terminal device generating the inference task—either an NPU or a TPU. In this case, the task execution plugin can directly determine the second target model from the third inference model. Furthermore, in this scenario, the task execution plugin can download the second target model to its own internal storage via the model management module, allowing the second target model to execute the inference task.
[0081] In this scenario, the third inference model can be obtained by converting the original inference model's format using the task execution plugin. The original inference model can be applied to common system-on-a-chip (SoC) devices such as CPUs and GPUs. Furthermore, the model's format conversion can be performed in advance and is independent of whether an inference task is generated. Specifically, the model's format conversion can be achieved using the conversion tool within the task execution plugin.
[0082] As can be seen, this approach involves pre-converting the format. After obtaining the inference task, the second target model is determined from the third inference model obtained after format conversion. Furthermore, in this case, the distribution of the second target model and the execution of the inference task can be combined... Figure 5 understand.
[0083] Alternatively, the system-on-a-chip (SoC) applicable to the third inference model may differ from the SoC provided by the terminal device generating the inference task. For example, the SoC applicable to the third inference model might be an NPU or TPU, while the SoC provided by the terminal device might be a CPU or GPU. In this case, in response to the acquisition of the inference task, if the task execution plugin can determine the fourth target model corresponding to the intelligent function within the third inference model, the task execution plugin can perform real-time format conversion of the fourth target model based on the type of SoC provided by the terminal device to obtain the second target model. If the SoC applicable to this second target model is the same as the SoC provided by the terminal, then this second target model can directly execute the inference task within the target runtime environment provided by the terminal device.
[0084] As can be seen, this situation is actually a way of converting the model format in real time in response to the acquisition of the inference task, so as to obtain a second target model for performing the inference task.
[0085] Regardless of whether the model format is pre-converted or converted in real-time, for terminal devices equipped with the task execution plugin, even if the system-on-chip (SoC) provided by the terminal device and the SoC applicable to the inference model executing the inference task are different, the model format conversion function provided by the task execution plugin can ensure that they are identical. Therefore, no matter what SoC the terminal device provides, the inference model can execute inference tasks on that SoC through format conversion. In other words, the format conversion function provided by the task execution plugin enables the task execution plugin to be compatible with the SoC of the terminal device. This compatibility ensures that inference tasks can be executed normally on fragmented SoCs, thus improving the problem of difficulty in implementing intelligent functions on terminal devices due to the fragmentation of SoCs.
[0086] It should also be noted that the above embodiments describe that the third inference model can be deployed in the task execution plugin. In the above case, "deployment" means that the task execution plugin can call the third inference model without using the network.
[0087] S303, if it is determined that a second target model exists based on the third inference model, then the second acquisition interface provided by the task execution plugin is called to obtain the execution pipeline of the inference task.
[0088] S304 runs the execution pipeline to perform inference tasks using a second target model running on a neural network processor.
[0089] After obtaining the second target model, the task execution plugin can load it onto the NPU. Simultaneously, the plugin can call its own second acquisition interface to obtain the execution pipeline for the inference task. This pipeline describes the specific process of the inference model performing the inference task, such as vectorization and pooling. The task execution plugin can then run the pipeline to perform the inference task using the second target model running on the NPU.
[0090] Optionally, the task execution plugin can also run an execution pipeline to accelerate the operation of the NPU, thereby improving the execution efficiency of inference tasks.
[0091] Optionally, to ensure the normal execution of the inference task, after retrieving the second target model, the task execution plugin can preprocess the input data of the inference task according to the configuration file of the second target model, and input the preprocessed results into the second target model so that the second target model can execute the inference task. Optionally, the preprocessing of the input data may include cropping, vectorizing, etc.
[0092] In this embodiment, for inference models that are not applicable to the NPU, if the model is assigned to an inference task, the model format conversion function of the task execution plugin can be used to enable the inference model to run on the NPU of the terminal device. This ensures the normal execution of the inference task and also improves the execution speed of the inference model.
[0093] On the other hand, for terminal devices equipped with task execution plugins, regardless of the system-on-a-chip (SoC) provided by the terminal device or the SoC to which the inference model is applicable, the model format conversion function provided by the task execution plugin can ensure that the SoC provided by the terminal device and the SoC applicable to the inference model executing the inference task are the same. Therefore, the inference model can execute the inference task on the SoC. In other words, the format conversion function provided by the task execution plugin ensures compatibility between the task execution plugin and the SoC of the terminal device. This compatibility guarantees that the inference task can be executed normally on fragmented SoCs, thus improving the problem of difficulty in implementing intelligent functions on terminal devices due to the fragmentation of SoCs.
[0094] Based on the above Figures 1-3 The method embodiment shown can be further described below from the perspective of internal structure, illustrating the working process of the task execution plugin. Therefore, the task execution plugin provided in this application embodiment can be described as follows: The terminal device can generate an inference task in response to the user's operation of the intelligent function. The task execution plugin can then acquire this inference task. Subsequently, the task execution plugin can utilize its capability detection function to determine the target task execution unit from its provided candidate task execution units, and finally execute the inference task using the target task execution unit within the target runtime environment provided by the terminal device. The capability detection function of the task execution plugin is essentially a task scheduling function, which can be used to schedule the inference task to the target task execution unit.
[0095] For details and examples regarding intelligent functions, reasoning tasks, and alternative task units, please refer to [link / reference]. Figure 1 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0096] In this embodiment, since the task execution plugin has built-in task execution units suitable for different operating environments, i.e., alternative task execution units, the plugin can be considered a plugin compatible with multiple operating environments. Due to this compatibility, after the plugin is installed on the terminal device and an inference task is generated, regardless of the terminal device's operating environment, the task execution plugin can determine the appropriate task execution unit from the alternative task execution units, and this task execution unit ensures the normal execution of the inference task.
[0097] In addition, the compatibility of the task execution plugin can ensure that the inference task can be executed normally in fragmented operating environments (i.e., multiple operating environments). At the same time, it can also reduce the difficulty of implementing intelligent functions on terminal devices. That is, there is no need to develop corresponding modules for executing inference tasks for different operating environments, and the terminal device can provide rich intelligent functions.
[0098] Figure 6 Another task execution plugin provided in the embodiments of the present invention, such as Figure 6 As shown, optionally, the plugin may include a task scheduling unit. This task scheduling unit can determine a target task execution unit from among the candidate task execution units and schedule the inference task to the target task execution unit. Furthermore, different task units among the candidate task execution units can correspond to different intelligent functions on the terminal device.
[0099] Optionally, such as Figure 6 As shown, the plugin may also include a first execution module and a second execution module.
[0100] Optionally, the alternative task execution unit may include an alternative inference interface, which has the highest query priority. Furthermore, the alternative inference interface can be integrated into the first execution module of the task execution plugin. The alternative inference interface in the alternative task execution unit can be used to call execution components deployed in the terminal device, and these execution components may specifically include a first inference model or rule algorithm with inference task execution capabilities.
[0101] Optionally, the alternative task execution unit may further include a second inference model deployed locally on the terminal device. This second inference model can be invoked by the second execution module in the task execution plugin. The invoked inference model is used to execute the inference task, and the number of parameters in the second inference model is greater than that of the first inference model. Furthermore, the inference tasks executed by the execution component invoked by the alternative inference interface and the inference tasks executed by the second inference model can correspond to different intelligent functions.
[0102] The collaborative working process between the first execution module, the second execution module, and the task scheduling unit in the task execution plugin can be described as follows: After the task scheduling unit obtains the reasoning task, since one type of reasoning task belongs to one intelligent function, it can determine whether there is a target reasoning interface corresponding to the intelligent function among the candidate reasoning interfaces integrated in the first execution module, based on the type of reasoning task.
[0103] In one scenario, if a target inference interface exists among the candidate inference interfaces, the task scheduling unit can schedule the inference task to the target inference interface. The execution component invoked by the target inference interface can then directly execute the inference task within the target runtime environment. In another scenario, if a target inference interface does not exist among the candidate inference interfaces, the task execution plugin can continue scheduling inference tasks. That is, it continues to determine, based on the type of inference task, whether a first target model corresponding to the intelligent function exists within the second inference model.
[0104] In one scenario, if the first target model exists within the second inference model, the task execution plugin can schedule the inference task to the second execution unit, which then calls the first target model, and the first target model directly executes the inference task in the target runtime environment. In another scenario, if the first target model does not exist within the second inference model—that is, if all candidate task execution units in the task execution plugin have been queried but no target task execution unit has been found—it indicates that the terminal device cannot execute the inference task. In this case, the task execution plugin can output a notification message reflecting the task execution failure, which can also be displayed to the user via the terminal device.
[0105] In this embodiment, the task execution plugin can perform queries sequentially according to query priority, using the task execution module to ensure the normal execution of the inference task by utilizing the queried target task execution units. Furthermore, the multi-level query priority setting improves search speed while reducing search costs. The two-level task execution unit setup also enriches the intelligent functions provided by the terminal device.
[0106] On the other hand, since different alternative inference interfaces in the alternative task execution unit correspond to different operating systems (i.e., fragmented operating systems), the compatibility of the task execution component can be specifically reflected in the compatibility with the terminal device operating system. This compatibility ensures that the inference task can be executed normally in fragmented operating systems, which improves the problem that terminal devices have difficulty implementing intelligent functions due to the fragmentation of terminal device operating systems.
[0107] like Figure 6As shown, optionally, the task execution plugin may further include a third execution module, and the alternative task execution unit may further include a third inference model. The number of parameters in the third inference model may be between that of the first and second inference models. The query priority of the third inference model is lower than that of the first inference model but higher than that of the second inference model. Furthermore, the third inference model can be deployed within the task execution plugin and can be called by the third execution module within the task execution plugin.
[0108] The collaborative working process between the first execution module, the second execution module, the third execution module, and the task scheduling unit in the task execution plugin can be described as follows: After the task scheduling unit obtains the reasoning task, it can determine whether there is a target reasoning interface corresponding to the intelligent function from the candidate reasoning interfaces integrated in the first execution module, based on the type of reasoning task.
[0109] In one scenario, if a target inference interface exists among the candidate inference interfaces, the task scheduling unit can schedule the inference task to the target inference interface. The execution component invoked by the target inference interface can then directly execute the inference task within the target runtime environment. In another scenario, if a target inference interface does not exist among the candidate inference interfaces, the task execution plugin can continue scheduling the inference task, that is, it continues to determine whether a second target model corresponding to the intelligent function exists based on the third inference model.
[0110] In one scenario, if a second target model exists, the task execution plugin can schedule the inference task to a third execution unit, which then calls the second target model, and the second target model directly executes the inference task within the target runtime environment. In another scenario, if a second target model does not exist, the task execution plugin can continue scheduling the inference task, that is, continue to determine within the second inference model whether a first target model corresponding to the intelligent function exists.
[0111] In one scenario, if the first target model exists within the second inference model, the task execution plugin can schedule the inference task to the second execution unit, which then calls the first target model, and the first target model directly executes the inference task in the target runtime environment. In another scenario, if the first target model does not exist within the second inference model—that is, if all candidate task execution units in the task execution plugin have been queried but no target task execution unit has been found—the task execution plugin can output a notification message indicating task execution failure. This notification message can also be displayed to the user via a terminal device.
[0112] This embodiment features a three-level query priority system. Similar to the embodiments described above, the task execution plugin can perform queries sequentially according to the query priority, leveraging the task execution module. This multi-level query priority setting improves search speed while reducing search costs, ensuring the normal execution of the inference task. Furthermore, the multi-level task execution unit configuration enriches the intelligent functions offered by the terminal device. On the other hand, the compatibility of the task execution component is specifically reflected in its compatibility with the terminal device's operating system. This compatibility ensures that the inference task can be executed normally even on fragmented operating systems.
[0113] Furthermore, the contents not described in detail in this embodiment and the technical effects that can be achieved can be found in the relevant descriptions in the above-mentioned related embodiments, and will not be repeated here.
[0114] like Figure 6 As shown, optionally, the plugin may also include a fourth execution module, and the alternative task execution unit may further include a fourth inference model. The fourth inference model can be deployed on a cloud server and can be called by the fourth execution module of the task execution plugin.
[0115] The collaborative working process between the first to fourth execution modules and the task scheduling unit in the task execution plugin can be described as follows: Once the reasoning task is obtained, if the task execution plugin fails to find a target task execution unit for executing the reasoning task during the three rounds of queries mentioned in the above embodiments, the task execution plugin can continue to schedule the reasoning task, that is, continue to determine whether there is a third target model corresponding to the intelligent function in the fourth reasoning model.
[0116] In one scenario, if a third target model exists within the fourth inference model, the task execution plugin can schedule the inference task to the fourth execution unit, which then invokes the third target model, which directly executes the inference task within the target runtime environment. In another scenario, if a third target model does not exist within the fourth inference model, the task execution plugin can output a notification message indicating task execution failure, which can also be displayed to the user via a terminal device.
[0117] This embodiment features a four-level query priority. Similar to the embodiments described above, the task execution plugin can perform queries sequentially according to the query priority, utilizing the task execution module. While ensuring the normal execution of the inference task, the multi-level query priority setting improves search speed and reduces search costs. Furthermore, the multi-level task execution unit configuration enriches the intelligent functions offered by the terminal device. On the other hand, the compatibility of the task execution component is specifically reflected in its compatibility with the terminal device's operating system. This compatibility ensures that the inference task can be executed normally even on fragmented operating systems.
[0118] Furthermore, the contents not described in detail in this embodiment and the technical effects that can be achieved can be found in the relevant descriptions in the above-mentioned related embodiments, and will not be repeated here.
[0119] like Figure 6 As shown, optionally, the plugin may also include an initialization unit.
[0120] The initialization unit can determine the first acquisition interface provided by the initialization unit that corresponds to the type of inference task, and then call the first acquisition interface to obtain the task initialization pipeline.
[0121] Subsequently, the task scheduling unit can run the task initialization pipeline to schedule the inference task, that is, to determine the target task execution unit from the candidate task execution units. At the same time, the task scheduling unit can also run the task initialization pipeline to create the execution environment for the inference task in the target runtime environment, so that the target task execution unit can execute the inference task in the execution environment.
[0122] It should be noted that the working process of the task initialization unit and the working process of the task scheduling unit provided in the above embodiments can be carried out synchronously or asynchronously. However, it is easy to understand that after the task scheduling unit schedules the inference task to the target task execution unit, the target task scheduling unit can only execute the inference task after the initialization is completed.
[0123] In this embodiment, the initialization unit determines the task initialization process for the inference task, which prepares for the formal execution of the inference task. The task scheduling unit can run this task initialization process to truly complete the preparations before the execution of the inference task, thereby ensuring the normal execution of the inference task.
[0124] Alternatively, the present application may be installed. Figure 6 The architecture diagram of the terminal device for the task execution plugin provided in the illustrated embodiment can also be combined with... Figure 7 Understanding. For example... Figure 7 As shown, the framework can be divided into the application layer, adaptation layer, service layer, kernel layer, and hardware abstraction layer from top to bottom.
[0125] The application layer can include at least one intelligent function provided by the terminal device. Taking an e-commerce app as an example, its intelligent functions can include personalized recommendations, voice interaction, image search, intelligent Q&A, etc. The adaptation layer can include the task scheduling unit and initialization unit in the task execution plugin. The service layer can include the first to fourth execution modules in the task execution plugin. The kernel layer can include the operating system provided by the terminal device, and the hardware abstraction layer can include the system-on-a-chip provided by the terminal device.
[0126] Based on the above Figures 4-5 The method embodiment shown, Figure 8 Another task execution plugin provided as an embodiment of the present invention may optionally include, for example: Figure 8 As shown, the third execution module may also include a conversion subunit.
[0127] This conversion subunit can convert the model format based on the type of system-on-a-chip in the target operating environment to obtain a model suitable for the NPU. This conversion process can be performed in advance without depending on the generation of the inference task, or it can be performed in real-time in response to the generation of the inference task. The specific conversion process can be referenced in [the documentation / reference]. Figure 4 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0128] In this embodiment, for inference models that are not applicable to the NPU, if the model is assigned to an inference task, the model format conversion function of the conversion subunit can be used to enable the inference model to run on the NPU of the terminal device so that the inference task can be executed normally. Executing the inference task on the NPU can also improve the execution speed of the inference model.
[0129] On the other hand, the format conversion function provided by the conversion subunit enables the task execution plug-in to be compatible with the system-on-a-chip of the terminal device. This compatibility can ensure that inference tasks can be executed normally on fragmented system-on-a-chips, which improves the problem of difficulty in implementing intelligent functions on terminal devices due to the fragmentation of system-on-a-chips.
[0130] Optionally, such as Figure 8 As shown, the third execution module in the plugin may include a loading subunit and an acceleration subunit. The loading subunit is equivalent to the inference engine, which can load the second target model onto the NPU provided by the terminal device. The process of determining the second target model can be found in the descriptions of the relevant embodiments above, and will not be repeated here. The acceleration subunit can accelerate the running of the second target model on the NPU, and the second target model running at an accelerated speed on the NPU can perform inference tasks.
[0131] In this embodiment, the use of the acceleration subunit can improve the execution speed of the inference model.
[0132] Optionally, such as Figure 8 As shown, the third execution module may further include a second acquisition interface. The third execution unit can call this second acquisition interface to obtain the execution pipeline for the inference task. In response to the execution pipeline, the second target model running on the NPU can perform the inference task.
[0133] Optionally, the third execution module may also include a run interface and a destroy interface, but these are not specified in the provided text. Figure 8 As shown in the diagram. The third execution unit can call this runtime interface to run the aforementioned execution pipeline. After the second target model completes its inference task, the third execution unit can call the destruction interface to destroy the aforementioned execution pipeline.
[0134] Optionally, such as Figure 8 As shown, the third execution module may further include a preprocessing subunit. This preprocessing subunit can preprocess the input data of the inference task according to the configuration file of the second target model, so that the second target model can use the preprocessed results as the input data for the inference task and execute the inference task.
[0135] In this embodiment, obtaining the execution pipeline and preprocessing the data both ensure the normal execution of the inference task. Furthermore... Figure 8 The third execution unit shown can also be represented architecturally as follows: Figure 9 As shown.
[0136] like Figure 9 As shown, the framework can be divided into an interface layer, a processing layer, and an engine layer from top to bottom.
[0137] The interface layer can include interfaces for creating, running, and destroying execution pipelines, and will handle... Figure 8 As described in the illustrated embodiment, the third execution module calls the second acquisition interface to obtain the execution pipeline of the inference task; subsequently, the third execution module can also call the run interface to run the execution pipeline. After the inference task is completed, the third execution module can also call the destroy interface to destroy the execution pipeline.
[0138] The processing layer may include preprocessing subunits and transformation subunits from the third execution module. Optionally, the processing layer may also store the execution pipeline obtained through calls from the interface layer, and may also include the respective tools required during the execution of the inference task, such as the model downloader, model parser, and patcher.
[0139] The engine layer can include loading subunits and acceleration subunits. Furthermore, these subunits within the engine layer can operate based on a system-on-a-chip (SoC) provided by the terminal device. In addition, for details not described in this embodiment and the technical effects that can be achieved, please refer to [link to relevant documentation]. Figure 4 and Figure 5 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0140] The task execution plugins provided in the above embodiments are all compatible. This compatibility can be reflected in compatibility with different operating systems provided by different terminal devices and different system-on-a-chips provided by different terminal devices. This compatibility can ensure that inference tasks can be executed normally in fragmented operating systems and / or fragmented system-on-a-chips.
[0141] Based on this, Figure 10 This is a flowchart illustrating another task execution method provided in this application. This method can be executed by a task execution plugin installed on a terminal device. However, it should be noted that, unlike the plugins mentioned in the above embodiments, the compatibility of this plugin lies in its compatibility with different system-on-a-chip (SoC) architectures. For example... Figure 10 As shown, the method may include the following steps: S401, Obtain the inference task generated using the intelligent functions on the terminal device.
[0142] The specific execution process of step S401 in this embodiment can be found in the description of the above-mentioned related embodiments, and will not be repeated here.
[0143] S402, based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is converted into a format to obtain the target inference model.
[0144] S403 performs inference tasks using a target inference model running on a neural network processor provided by the terminal device.
[0145] When the system-on-a-chip (SoC) applicable to the candidate inference model provided by the task execution plugin differs from that provided by the terminal device—for example, if the candidate inference model is applicable to an NPU or TPU, while the terminal device provides a CPU or GPU—the task execution plugin can convert the format of the candidate inference model according to the type of SoC provided by the terminal device to obtain the target inference model. Ultimately, the task execution plugin can then use the target inference model running on the NPU to perform inference tasks.
[0146] It should be noted that the target inference model in this embodiment is the same as the second target model mentioned in the above embodiments. Furthermore, regarding the conversion of the model format, and... Figure 4 The illustrated embodiment also has two conversion methods: pre-conversion and real-time conversion. For details of the conversion process, please refer to [link / reference]. Figure 4 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0147] When performing a pre-conversion, the candidate reasoning model mentioned in this embodiment is the same as the third reasoning module mentioned in the above embodiments. When performing a real-time conversion, the candidate reasoning model mentioned in this embodiment is the same as the fourth target model mentioned in the above embodiments, which is selected from the third reasoning module and corresponds to the intelligent function.
[0148] In this embodiment, regardless of whether the model format is pre-converted or converted in real-time, for a terminal device equipped with a task execution plugin, even if the system-on-a-chip (SoC) provided by the terminal device and the SoC applicable to the inference model executing the inference task are different, the model format conversion function provided by the task execution plugin can ensure that they are identical. Therefore, regardless of the SoC provided by the terminal device, the inference model can execute inference tasks on this SoC through format conversion. In other words, the format conversion function provided by the task execution plugin enables the task execution plugin to be compatible with the SoC of the terminal device. This compatibility ensures that the inference task can be executed normally on fragmented SoCs, thus improving the problem of difficulty in implementing intelligent functions on terminal devices due to the fragmentation of SoCs.
[0149] Based on the above Figure 10 The method embodiment shown can be further described below from the perspective of internal structure, illustrating the working process of the task execution plugin. Therefore, the task execution plugin provided in this application embodiment can be described as follows: The task execution plugin can acquire inference tasks generated using the intelligent features on the terminal device. Then, Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is converted into a new format to obtain the target inference model. The inference task is performed using a target inference model running on a neural network processor provided by the terminal device.
[0150] Optionally, such as Figure 11 As shown, Figure 10 The task execution plugin corresponding to the embodiment may further include: a conversion subunit, used to convert the format of the candidate model to obtain the target inference model.
[0151] Optionally, such as Figure 11 As shown, Figure 10The task execution plugin corresponding to the embodiment may further include: a preprocessing subunit, used to preprocess the input data of the reasoning task according to the configuration file of the target reasoning model, so that the target reasoning model can use the preprocessing result as the input data of the reasoning task to execute the reasoning task.
[0152] Optionally, Figure 10 The task execution plugin corresponding to the embodiment may further include: an acquisition interface, a run interface, and a destroy interface. These interfaces are in Figure 11 It is not shown in the middle.
[0153] The task execution plugin can call the retrieval interface to obtain the execution pipeline corresponding to the inference task. Then, the task execution plugin can also call the run interface to run the aforementioned execution pipeline. After the second target model completes the inference task, the task execution plugin can call the destruction interface to destroy the aforementioned execution pipeline.
[0154] Optionally, such as Figure 11 As shown, Figure 10 The task execution plugin corresponding to the embodiment may also include: a loading subunit and an acceleration subunit.
[0155] The loading subunit is used to load the target inference model onto the neural network processor.
[0156] The acceleration subunit is used to accelerate the execution of the target inference model on the neural network processor, so that the accelerated target inference model can perform inference tasks.
[0157] The specific working process of each subunit in this embodiment can be found in [reference]. Figure 8 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0158] For details not described in this embodiment and the technical effects that can be achieved, please refer to the above. Figure 4 , Figure 8 and Figure 9 The descriptions in the embodiments will not be repeated here.
[0159] It should be noted that, although the executing entity of each step is specified in the description of the steps provided in the above embodiments, this application does not limit the steps in the above method embodiments to the same device. That is, each step in the method embodiment can be executed by the same device, or the method can be executed by different devices. For example, the execution subject of steps S101 to S103 can be device A; or the execution subject of steps S101 and S102 can be device A, and the execution subject of step S103 can be device B; and so on.
[0160] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as S101, S102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0161] The cloud network management system and performance monitoring engine provided in the above embodiments of this application can be deployed on server-side devices in a cloud computing environment. Users can trigger tasks to the server-side devices in the system via client devices. Optionally, the server-side devices can be cloud servers maintained by cloud service providers—referred to as nodes. Client devices can be laptops, tablets, PCs, robots, etc.
[0162] In such Figure 12 The cloud computing environment shown may include several distributed deployments. Figure 12 The diagram illustrates compute nodes (201-1, 201-2, ...), and cache nodes. Each node can possess processing resources such as computing and storage. In a cloud computing environment, multiple nodes can be organized to provide a certain service, such as the cloud services mentioned in the embodiments of this application, including computing services, storage services, and database services. Of course, a single node can also provide one or more services, such as... Figure 12 The diagram illustrates services A, B, C, and D. In a cloud computing environment, services can be provided through external service interfaces, which client devices can call to use the corresponding services. Service interfaces include Software Development Kits (SDKs) and Application Programming Interfaces (APIs), among others.
[0163] The services described above are deployed using various virtualization technologies supported by cloud computing environments, such as virtual machine-based and container-based virtualization technologies. Taking container-based virtualization technology as an example, several containers corresponding to a service can be assembled into a container group (pod). For example... Figure 12The illustrated service B can be configured with one or more pods, and each pod can include a proxy and one or more containers. The one or more containers in the pod are used to handle requests related to one or more corresponding functions of the service, and the proxy in the pod is used to control network functions related to the service, such as routing and load balancing.
[0164] During operation, executing requests from client devices may require invoking one or more services in the cloud computing environment, and executing one or more functions of one service may require invoking one or more functions of another service. For example... Figure 12 As shown, after receiving a request from a client device, service A can call service B, and service B can request service D to perform one or more functions.
[0165] In one possible design, the task execution plugins provided in the above embodiments can also be applied to an electronic device. For example... Figure 13 As shown, the electronic device may include a processor 21 and a memory 22. The memory 22 stores programs that support the electronic device in performing various operations that the task execution plugin described above can perform. The memory 22 stores programs that support the electronic device in performing the aforementioned operations. Figures 1-4 The program for the task execution method provided in the illustrated embodiment is such that the processor 21 is configured to execute the program stored in the memory 22.
[0166] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the processor 21, they can perform the following steps: Acquire inference tasks generated using the intelligent functions on the terminal device; Among the alternative task execution units provided by the task execution plugin built into the terminal device, the target task execution unit is determined. The different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device. The inference task is executed using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task.
[0167] Optionally, different task units among the alternative task execution units correspond to different intelligent functions on the terminal device; The alternative task execution unit includes alternative inference interfaces. Different inference interfaces in the alternative inference interfaces correspond to different operating systems of the terminal device. The alternative inference interfaces are used to call the execution components deployed in the terminal device. The execution components include a first inference model and / or rule algorithm with inference task execution capability. The alternative task execution unit further includes a second inference model, which has a larger number of parameters than the first inference model.
[0168] Optionally, when determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically used to: determine whether there is a target inference interface corresponding to the intelligent function among the alternative inference interfaces.
[0169] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically configured to: if the target inference interface exists among the alternative inference interfaces, then execute the inference task using the execution component called by the target inference interface in the target operating environment.
[0170] Optionally, when determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically configured to: if the target inference interface does not exist among the alternative inference interfaces, then determine whether there is a first target model corresponding to the intelligent function in the second inference model.
[0171] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically configured to: if the first target model exists in the second inference model, then execute the inference task using the first target model in the target operating environment.
[0172] Optionally, the alternative task execution unit further includes a third inference model, wherein the number of parameters of the third inference model is between that of the first inference model and the second inference model; When determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically used to: if the target inference interface does not exist among the alternative inference interfaces, then determine whether there is a second target model corresponding to the intelligent function according to the third inference model.
[0173] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically configured to: if the second target model exists, then execute the inference task using the second target model in the target operating environment.
[0174] Optionally, when determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically used to: if the second target model does not exist, then determine whether a first target model corresponding to the intelligent function exists in the second inference model.
[0175] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically configured to: if the first target model exists in the second inference model, then execute the inference task using the first target model in the target operating environment.
[0176] Optionally, the alternative task execution unit may further include a fourth inference model deployed in a cloud server.
[0177] When determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically used to: if the first target model does not exist in the second inference model, then determine whether a third target model corresponding to the intelligent function exists in the fourth inference model.
[0178] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically used to: if the third target model exists in the fourth inference model, then execute the inference task using the third target model in the target operating environment.
[0179] Optionally, after acquiring the inference task generated by using the intelligent function on the terminal device, the processor 21 is further configured to: determine the first acquisition interface provided by the task execution plugin according to the type of the inference task; and call the first acquisition interface to acquire the task initialization pipeline corresponding to the inference task.
[0180] When determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device, the processor 21 is specifically configured to: run the task initialization pipeline to determine the target task execution unit from the alternative task execution units.
[0181] Optionally, the processor 21 is further configured to run the task initialization pipeline to create the execution environment for the inference task in the target runtime environment.
[0182] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically used to: execute the inference task using the target task execution unit in the execution environment.
[0183] The system-on-a-chip in the target operating environment includes a neural network processor.
[0184] Optionally, the processor 21 is further configured to, if it is determined according to the third inference model that the second target model exists, invoke the second acquisition interface provided by the task execution plugin to obtain the execution pipeline of the inference task.
[0185] When the processor 21 executes the inference task using the target task execution unit in the target operating environment provided by the terminal device that generates the inference task, it is specifically used to: run the execution pipeline to execute the inference task using the second target model running on the neural network processor.
[0186] Optionally, when the processor 21 runs the execution pipeline to perform the inference task using the second target model running on the neural network processor, it is specifically configured to: load the second target model onto the neural network processor; and run the execution pipeline to perform the inference task using the second target model running at an accelerated speed on the neural network processor.
[0187] Optionally, the format of the third inference model is not applicable to the neural network processor.
[0188] The processor 21 is further configured to, if there is a fourth target model in the third inference model that corresponds to the intelligent function, convert the format of the fourth target model according to the type of system-on-a-chip in the target operating environment to obtain the second target model.
[0189] Optionally, the format of the third inference model is suitable for the neural network processor; the processor 21 is further configured to determine the second target model from the third inference model according to the model management module.
[0190] Optionally, the processor 21 is further configured to preprocess the input data of the inference task according to the configuration file of the second target model; When the processor 21 performs the inference task using the second target model, it is specifically used to: input the preprocessing result into the second target model so that the second target model can perform the inference task.
[0191] Optionally, Figure 13The memory 22 in the illustrated electronic device is also used to store programs that support the electronic device in performing various operations that the task execution plug-in described above can perform. The memory 22 is also used to store programs that support the electronic device in performing the aforementioned operations. Figure 10 The program for the task execution method provided in the illustrated embodiment is such that the processor 21 is configured to execute the program stored in the memory 22.
[0192] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the processor 21, they can perform the following steps: Acquire inference tasks generated using the intelligent functions on the terminal device; Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is format-converted to obtain the target inference model; The inference task is performed using the target inference model running on a neural network processor provided by the terminal device.
[0193] The structure of the electronic device may also include other components such as a communication component 23, a display 24, a power supply component 25, and an audio component 26.
[0194] Figure 13 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 13 The components shown.
[0195] in addition, Figure 13 The components mentioned are optional, not mandatory, and depend on the specific product form of the electronic device. The electronic device in this application embodiment can be a conventional server, cloud server, or server array, etc.
[0196] The processor 21 can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a CPU, a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), a complex programmable logic device (CPLD), or other programmable devices; or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited to these.
[0197] The aforementioned memory 22 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0198] The aforementioned communication component 23 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as Wi-Fi, 2G (e.g., Global System for Mobile Communications (GSM)), 3G (e.g., Wideband Code Division Multiple Access (WCDMA), 4G (e.g., Long Term Evolution (LTE)), 4G+ (e.g., LTE-Advanced (LTE-A)), or 5G (5th Generation Mobile Communication Technology), or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), Infrared Data Association (IRDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0199] The aforementioned display 24 includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0200] The aforementioned power supply component 25 provides power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0201] The audio component 26 described above can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0202] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), SRAM, dynamic random access memory (DRAM), other types of random-access memory (RAM), ROM, EEPROM, EPROM, PROM, flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.
[0203] Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. These computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device can be implemented as a means to implement the corresponding functions in the above method embodiments.
[0204] Furthermore, the specific implementation form of the computer program product is not limited in the embodiments of this application. In some embodiments, the computer program product may be implemented as an application (APP), a mini-program, a PC client, a program module, a plug-in, an installation package, a software development kit (SDK), an image file of an optical disc (such as an ISO file), a plug-in, or software in the form of Software as a Service (SaaS), etc., but is not limited to these.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A task execution method, characterized in that, The method includes: Acquire inference tasks generated using the intelligent functions on the terminal device; Among the alternative task execution units provided by the task execution plugin built into the terminal device, the target task execution unit is determined. The different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device. The inference task is executed using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task.
2. The method according to claim 1, characterized in that, The different task units in the alternative task execution units correspond to different intelligent functions on the terminal device; The alternative task execution unit includes alternative inference interfaces. Different inference interfaces in the alternative inference interfaces correspond to different operating systems of the terminal device. The alternative inference interfaces are used to call the execution components deployed in the terminal device. The execution components include a first inference model and / or rule algorithm with inference task execution capability. The alternative task execution unit further includes a second inference model, which has a larger number of parameters than the first inference model.
3. The method according to claim 2, characterized in that, The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: Among the candidate reasoning interfaces, determine whether there exists a target reasoning interface corresponding to the intelligent function; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: If the target inference interface exists among the candidate inference interfaces, then in the target runtime environment, the inference task is executed using the execution component called by the target inference interface.
4. The method according to claim 3, characterized in that, The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: If the target reasoning interface is not among the candidate reasoning interfaces, then in the second reasoning model, it is determined whether there is a first target model corresponding to the intelligent function; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: If the first target model exists in the second inference model, then the inference task is executed using the first target model in the target operating environment.
5. The method according to claim 3, characterized in that, The alternative task execution unit further includes a third inference model, the number of parameters of which is between that of the first inference model and the second inference model; The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: If the target reasoning interface is not among the candidate reasoning interfaces, then based on the third reasoning model, it is determined whether there is a second target model corresponding to the intelligent function; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: If the second target model exists, then the reasoning task is performed using the second target model in the target operating environment.
6. The method according to claim 5, characterized in that, The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: If the second target model does not exist, then in the second reasoning model, determine whether there exists a first target model corresponding to the intelligent function; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: If the first target model exists in the second inference model, then the inference task is executed using the first target model in the target operating environment.
7. The method according to claim 4 or 6, characterized in that, The alternative task execution unit also includes a fourth inference model deployed on a cloud server; The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: If the first target model is not present in the second reasoning model, then it is determined in the fourth reasoning model whether a third target model corresponding to the intelligent function exists. The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: If the third target model exists in the fourth inference model, then the inference task is executed using the third target model in the target operating environment.
8. The method according to claim 1, characterized in that, After obtaining the inference task generated using the intelligent functions on the terminal device, the method further includes: The first acquisition interface provided by the task execution plugin is determined based on the type of the inference task. Call the first acquisition interface to obtain the task initialization pipeline corresponding to the inference task; The process of determining the target task execution unit from the alternative task execution units provided by the task execution plugin built into the terminal device includes: Run the task initialization pipeline to determine the target task execution unit from the candidate task execution units.
9. The method according to claim 8, characterized in that, The method further includes: Run the task initialization pipeline to create the execution environment for the inference task in the target runtime environment; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: In the execution environment, the inference task is executed using the target task execution unit.
10. The method according to claim 5, characterized in that, The system-on-a-chip in the target operating environment includes a neural network processor; the method further includes: If the existence of the second target model is determined according to the third inference model, then the second acquisition interface provided by the task execution plugin is called to obtain the execution pipeline of the inference task; The step of executing the inference task using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task includes: The execution pipeline is run to perform the inference task using the second target model running on the neural network processor.
11. The method according to claim 10, characterized in that, Running the execution pipeline to perform the inference task using the second target model running on the neural network processor includes: Load the second target model onto the neural network processor; The execution pipeline is run to perform the inference task using the second target model, which is running at an accelerated speed on the neural network processor.
12. The method according to claim 10, characterized in that, The format of the third inference model is not applicable to the neural network processor; the method further includes: If a fourth target model corresponding to the intelligent function exists in the third inference model, then the fourth target model is format-converted according to the type of system-on-a-chip in the target operating environment to obtain the second target model.
13. The method according to claim 10, characterized in that, The format of the third inference model is suitable for the neural network processor; the method further includes: The second target model is determined from the third inference model according to the model management module.
14. The method according to claim 5, characterized in that, The method further includes: According to the configuration file of the second target model, the input data of the inference task is preprocessed; The step of performing the inference task using the second target model includes: The preprocessing results are input into the second target model so that the second target model can perform the inference task.
15. A task execution plugin, characterized in that, The plugin is configured to perform the following operations: Acquire inference tasks generated using the intelligent functions on the terminal device; Among the alternative task execution units provided by the task execution plugin built into the terminal device, the target task execution unit is determined. The different task execution units among the alternative task execution units are suitable for different operating environments, including the operating system and / or system-on-a-chip of the terminal device. The inference task is executed using the target task execution unit within the target operating environment provided by the terminal device that generates the inference task.
16. The plug-in according to claim 15, characterized in that, The different task units in the alternative task execution units correspond to different intelligent functions on the terminal device; The alternative task execution unit includes an alternative inference interface, which is integrated into the first execution module of the task execution plugin. The alternative inference interface is used to call the execution component deployed in the terminal device. The execution component includes a first inference model or rule algorithm with inference task execution capability. The plugin includes a task scheduling unit, used to determine whether there is a target inference interface corresponding to the intelligent function among the alternative inference interfaces; If the target inference interface exists among the candidate inference interfaces, the inference task is scheduled to the target inference interface so that the inference task can be executed in the target runtime environment using the execution component called by the target inference interface.
17. The plug-in according to claim 16, characterized in that, The alternative task execution unit includes a second inference model, the second inference model having more parameters than the first inference model, and the second inference model being called by the second execution module of the task execution plugin; The task scheduling unit is used to determine whether there is a first target model corresponding to the intelligent function in the second inference model if the target inference interface is not among the candidate inference interfaces. If the first target model exists in the second inference model, the inference task is scheduled to the first target model so that the inference task can be executed using the first target model in the target runtime environment.
18. The plug-in according to claim 16, characterized in that, The alternative task execution unit also includes a third inference model, the number of parameters of which is between that of the first inference model and the second inference model, and the third inference model is called by the third execution module of the task execution plugin; The plugin includes a task scheduling unit, which is used to determine whether there is a second target model corresponding to the intelligent function if the target inference interface is not among the candidate inference interfaces, based on the third inference model. If the second target model exists, the inference task is scheduled to the second target model so that the inference task can be executed using the second target model in the target runtime environment; If the second target model does not exist, then in the second reasoning model, determine whether there exists a first target model corresponding to the intelligent function; If the first target model exists in the second inference model, the inference task is scheduled to the first target model so that the inference task can be executed using the first target model in the target runtime environment.
19. The plug-in according to claim 17 or 18, characterized in that, The alternative task execution unit also includes a fourth inference model deployed in a cloud server, which is called by the fourth execution module of the task execution plugin; The task scheduling unit is used to determine whether a third target model corresponding to the intelligent function exists in the fourth inference model if the first target model does not exist in the second inference model. If the third target model exists in the fourth inference model, the inference task is scheduled to the third target model so that the inference task can be executed using the third target model in the target operating environment.
20. The plug-in according to claim 15, characterized in that, The plugin also includes an initialization unit, used to determine a first acquisition interface provided by the initialization unit according to the type of the inference task; and to call the first acquisition interface to obtain the task initialization pipeline. The task scheduling unit is used to run the task initialization pipeline to determine the target task execution unit from the candidate task execution units; The task initialization pipeline is run to create an execution environment for the inference task in the target runtime environment, so that the target task execution unit can execute the inference task in the execution environment.
21. The plug-in according to claim 18, characterized in that, The third execution module includes a preprocessing subunit, used to preprocess the input data of the inference task according to the configuration file of the second target model, so that the second target model uses the preprocessing result as the input data of the inference task to execute the inference task.
22. The plug-in according to claim 18, characterized in that, The system-on-a-chip in the target operating environment includes a neural network processor; the third execution module includes a second acquisition interface; The third execution module is used to call the second acquisition interface to obtain the execution pipeline of the inference task, and in response to the operation of the execution pipeline, to execute the inference task using the second target model running on the neural network processor.
23. The plug-in according to claim 22, characterized in that, The third execution module further includes: a loading subunit and an acceleration subunit; The loading subunit is used to load the second target model onto the neural network processor; The acceleration subunit is used to accelerate the execution of the second target model on the neural network processor, so as to perform the inference task using the second target model that is accelerated on the neural network processor.
24. The plug-in according to claim 22, characterized in that, The format of the third inference model is not applicable to the neural network processor; The third execution module includes a conversion subunit, which is used to convert the format of the fourth target model according to the type of system-on-a-chip in the target operating environment to obtain the second target model if there is a fourth target model in the third inference model that corresponds to the intelligent function.
25. A method for performing a reasoning task, characterized in that, include: Acquire inference tasks generated using the intelligent functions on the terminal device; Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is format-converted to obtain the target inference model; The inference task is performed using the target inference model running on a neural network processor provided by the terminal device.
26. A task execution plugin, characterized in that, The plugin is configured to perform the following operations: Acquire inference tasks generated using the intelligent functions on the terminal device; Based on the type of system-on-a-chip provided by the terminal device that generates the inference task, the candidate inference model provided by the task execution plugin is format-converted to obtain the target inference model; The inference task is performed using the target inference model running on a neural network processor provided by the terminal device.
27. The plug-in according to claim 26, characterized in that, The plugin includes a conversion subunit, which is used to convert the format of the candidate inference model corresponding to the intelligent function provided by the task execution plugin according to the system-on-a-chip type provided by the terminal device, so as to obtain the target inference model suitable for the neural network processor.
28. The plug-in according to claim 26, characterized in that, The plugin includes a preprocessing subunit, which is used to preprocess the input data of the inference task according to the configuration file of the target inference model, so that the target inference model can use the preprocessing result as the input data of the inference task and execute the inference task.
29. The plug-in according to claim 26, characterized in that, The plugin includes a loading subunit and an acceleration subunit; The loading subunit is used to load the target inference model onto the neural network processor; The acceleration subunit is used to accelerate the execution of the target inference model on the neural network processor, so that the accelerated target inference model can perform the inference task.
30. An electronic device comprising: A memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor performs the reasoning task execution method as described in any one of claims 1 to 14, or the task execution method as described in claim 25.
31. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the reasoning task execution method as described in any one of claims 1 to 14, or the task execution method as described in claim 25.
32. A computer program product, characterized in that, The computer program product includes a computer program or instructions that enable the computer program or instructions to implement the task execution method of any one of claims 1 to 14, or the reasoning task execution method of claim 25.