Inference Framework, Inference Method, Device, Equipment and Storage Medium

By using plug-ins to define modules and basic functional modules in the inference framework, encapsulating and unifying the functional interfaces of inference equipment, the complex requirements for various device technologies during inference model deployment are solved, and the effect of simplifying development and improving scalability is achieved.

CN114936643BActive Publication Date: 2025-05-27BEIJING SENSETIME TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210546113.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-05-27
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

In the field of artificial intelligence, it is necessary to understand the equipment management, model reasoning, preprocessing and postprocessing technologies of various reasoning devices when deploying inference models, which increases the difficulty of development.

Method used

Provides a reasoning framework, which encapsulates the functions of the reasoning device into an internal plug-in interface through the plug-in definition module, and generates a unified external plug-in interface through the basic functional module to isolate hardware differences and support multiple inference devices.

Benefits of technology

It reduces the coupling between different inference equipment and inference frameworks, simplifies the development process, shortens the R&D cycle, saves R&D costs, and improves the scalability and application scope of the framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936643B_ABST
    Figure CN114936643B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an inference framework, an inference method, a device, equipment, and a storage medium. The inference framework includes a plugin definition module and a basic function module. Among them, the plugin definition module includes function plugins corresponding to each of multiple inference devices. The function plugins are used to encapsulate the inference process into an internal plugin interface. The basic function module is used to nest and combine the internal plugin interfaces corresponding to each inference device to generate corresponding external plugin interfaces. The external plugin interfaces are used to respond to an inference request based on a target inference device, and call the internal plugin interface corresponding to the target inference device to convert the data to be inferred into service data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of artificial intelligence technology, and in particular, to an inference framework, an inference method, a device, a device, and a storage medium. Background Art

[0002] Artificial Intelligence (AI) is a technical science that studies, develops, and extends theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. In the field of artificial intelligence, corresponding functions can be implemented by deploying inference models in different inference devices. However, during the deployment of inference models, it is necessary to understand the corresponding device management, model inference, pre-processing, and post-processing technologies for different inference devices, and then combine these technologies to achieve a complete inference process, which increases the development difficulty of artificial intelligence-related applications. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure provide at least an inference framework, an inference method, a device, a device, and a storage medium.

[0004] The technical solution of the embodiments of the present disclosure is implemented as follows:

[0005] On the one hand, an inference framework is provided in an embodiment of the present disclosure. The inference framework includes a plugin definition module and a basic function module, where: the plugin definition module includes function plugins corresponding to each of the multiple inference devices; the function plugins are used to encapsulate the inference process into internal plugin interfaces; the basic function module is used to nest and combine the internal plugin interfaces corresponding to each inference device to generate corresponding external plugin interfaces; the external plugin interfaces are used to respond to an inference request based on a target inference device, and call the internal plugin interface corresponding to the target inference device to convert the data to be inferred into service data.

[0006] In some embodiments, the internal plugin interface includes an internal device interface, the external plugin interface includes an external device interface, and the function plugin includes a device plugin for providing the internal device interface. Among them, the basic function module nests and combines the internal device interfaces to generate an external device interface; when the inference request includes a device management request for the target inference device, the external device interface is used to call the internal device interface corresponding to the target inference device based on the device management request to manage the target inference device.

[0007] Based on the above embodiments, since the external device interface is provided externally through the basic function module, and this external device interface can call the internal device interfaces of multiple inference devices provided by the device plug-ins, thus, the hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external device interface is provided externally, thereby providing support for the large-scale production of the software industry, shortening the R & D cycle and saving the R & D cost. At the same time, since the internal device interfaces of each inference device are provided by the device plug-ins included in the plug-in definition module, when a new device is connected, only the device plug-in corresponding to this inference device needs to be defined, that is, the internal device interface corresponding to this inference device is defined, and the management of this new device can be supported. And since what is changed is the internal device interface, for the framework users, there is no need to change the upper-layer code of the framework, further saving the development cost and expanding the application scope of the framework.

[0008] In some embodiments, the internal device interface includes a device binding sub-interface, a device status acquisition sub-interface, and a device memory operation sub-interface; wherein: the device binding sub-interface is used to acquire the identification information of the inference device; the device status acquisition sub-interface is used to acquire the device status of the inference device; the device memory operation sub-interface is used to operate the device memory.

[0009] Based on the above embodiments, since the device plug-in can provide sub-interfaces with different functions for the inference device, during the process of defining the interfaces for device management for this inference device, the various functions of the inference device management can be separated, effectively reducing the coupling degree between different functional sub-interfaces and ensuring the stability of device management during the inference process. At the same time, during the process of using the inference device to process the inference task, if an inference error occurs, the problem location can be quickly located, and the developer only needs to re-define the sub-interface where the current problem appears, reducing the development cost.

[0010] In some embodiments, the internal plug-in interface includes an internal processing interface, the external plug-in interface includes an external processing interface, and the function plug-in includes a processing plug-in for providing the internal processing interface; wherein, the basic function module nests and combines the internal processing interfaces to generate the external processing interface; in the case where the inference request includes a pre-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the pre-processing request to convert the data to be inferred into the input data of the inference model; in the case where the inference request includes a post-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the post-processing request to convert the output data of the inference model into business data.

[0011] Based on the above embodiments, since an external processing interface is provided to the outside through the basic function module, and this external processing interface can call the internal processing interfaces of multiple inference devices provided by the processing plug-ins. Therefore, hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external processing interface is provided to the outside, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal processing interfaces of each inference device are provided by the processing plug-ins included in the plug-in definition module, when a new device is connected, only the processing plug-in corresponding to the inference device needs to be defined, that is, the internal processing interface corresponding to the inference device is defined, and the data processing of the inference process can be completed using this inference device, improving the scalability and deployment efficiency of the inference framework.

[0012] In some embodiments, the internal processing interface includes a pre-processing sub-interface and a post-processing sub-interface; the pre-processing sub-interface is used to convert the data to be inferred into the input data of the inference model; the post-processing sub-interface is used to convert the output data of the inference model into service data.

[0013] Based on the above embodiments, since the processing plug-in can provide sub-interfaces with different functions for the inference device, the pre-processing process and the post-processing process in the inference process can be separated during the process of defining the interfaces for the processing of the inference device, effectively reducing the coupling degree between different functional sub-interfaces and ensuring the stability of the data processing process during the inference process. At the same time, during the process of using the inference device to process the inference task, if an inference error occurs, the problem location can be quickly located. During the process of eliminating the problem, the developer only needs to re-define the sub-interface where the current problem occurs, reducing the development cost.

[0014] In some embodiments, the internal plug-in interface includes an internal inference interface, the external plug-in interface includes an external inference interface, and the functional plug-in includes an inference plug-in for providing the internal inference interface. Among them, the basic function module nests and combines the internal inference interfaces to generate an external inference interface. When the inference request includes a model inference request for the target inference device, the external inference interface is used to call the internal inference interface corresponding to the target inference device based on the model inference request, and convert the input data into output data based on the inference model.

[0015] Based on the above embodiments, since the external inference interface is provided externally through the basic function module, and this external inference interface can call the internal inference interfaces of multiple inference devices provided by the inference plug-ins. Thus, the hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external inference interface is provided externally, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal inference interfaces of each inference device are provided by the inference plug-ins included in the plug-in definition module, when a new device is connected, only the inference plug-in corresponding to this inference device needs to be defined, that is, the internal inference interface corresponding to this inference device needs to be defined, and then the inference process can be executed using this inference device, improving the scalability and deployment efficiency of the inference framework.

[0016] In some embodiments, the internal inference interface includes the internal inference interfaces corresponding to at least one inference model. The external inference interface is used to call the target internal inference interface corresponding to the target inference device based on the model inference request for the target inference device, and convert the input data into output data based on the target inference model. The inference request is used to call the target inference model in the at least one inference model, and the target inference interface is used to manage the inference process of the target inference model.

[0017] Based on the above embodiments, since the external processing interface is provided externally through the basic function module, and this external processing interface can call the internal processing interfaces corresponding to at least one inference model provided by the processing plug-ins. Thus, the model differences can be isolated at the framework level. While supporting multiple inference devices and inference models, a unified external inference interface is provided externally, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal inference interfaces corresponding to each inference model are provided by the inference plug-ins included in the plug-in definition module, in the case of needing to adapt to a new inference model, only the inference plug-in corresponding to this new inference model needs to be defined, that is, the internal processing interface corresponding to this new inference model needs to be defined, and then the input data can be converted into output data based on this new inference model, improving the scalability and deployment efficiency of the inference framework.

[0018] In some embodiments, the internal inference interface includes: a model parsing sub-interface, a model configuration sub-interface, a model inference sub-interface, and a result acquisition sub-interface. Among them, the model parsing sub-interface is used to obtain the model file and perform parsing. The model configuration sub-interface is used to configure the parsed model to obtain the configured model. The model inference sub-interface is used to complete the model inference process based on the configured model. The result acquisition sub-interface is used to obtain the inference result of the model inference process.

[0019] Based on the above embodiments, since the inference plugin can provide sub-interfaces with different functions for the inference device, during the process of defining the interface for the inference process of the inference device, the processes of model parsing, model configuration, model inference, and obtaining the interface in the inference process can be separated, effectively reducing the coupling degree between different functional sub-interfaces and ensuring the stability of the inference process based on the inference model. At the same time, during the process of using the inference device to process inference tasks, if an inference error occurs, the problem location can be quickly located. During the process of eliminating the problem, developers only need to redefine the sub-interface where the current problem occurs, reducing the development cost.

[0020] In some embodiments, the basic function module is used to provide at least one business function; when the business function is called, it performs at least one of the following: when the business function is called by the internal inference interface, converting the input data into the output data; when the business function is called by the internal processing interface, converting the data to be inferred into the input data of the inference model; when the business function is called by the internal processing interface, converting the output data of the inference model into business data.

[0021] Based on the above embodiments, since the business functions provided by the functional functions can enable each internal function plugin to reuse the business function, while improving the code reuse rate, it can also reduce the development difficulty.

[0022] In some embodiments, the basic function module is used to provide an exchange data structure corresponding to each inference device; the inference request is used to convert the data to be inferred in the first data structure into the business data in the second data structure; when the internal processing interface is called, it is also used to convert the data to be inferred in the first data structure into the data to be inferred in the third data structure; when the internal processing interface is called, it is also used to convert the business data in the fourth data structure into the business data in the second data structure; the third data structure and the fourth data structure are data structures supported by the target inference device.

[0023] Based on the above embodiments, since the basic function module provides an exchange data interface corresponding to each inference device, in the process of using different inference devices for data processing, the framework user can input the data to be inferred (the first data structure) with a unified data structure, and the inference framework can convert the data to be inferred into a data structure supported by the inference device (the third data structure) based on the actually deployed inference device; at the same time, after the inference framework outputs the service data of the fourth data structure, it can also be converted into the service data of the second data structure; thus, for the framework user, based on the unified upper-layer code, that is, using the unified data structure, the inference framework can convert the unified data structure into an exchange data structure supported by different inference devices through the basic function module, improving the parallelism of software development.

[0024] In some embodiments, the basic function module is used to provide a model data structure corresponding to each of the inference models; when the internal inference interface is called, it is also used to convert the input data of the fifth data structure into the input data of the sixth data structure; the internal inference interface is used to convert the input data of the sixth data structure into the output data of the seventh data structure based on the target inference model; when the internal inference interface is called, it is also used to convert the output data of the seventh data structure into the service data of the eighth data structure.

[0025] Based on the above embodiments, since the basic function module provides a model data structure corresponding to each of the inference models, in the process of using different inference models for model inference, the framework user can input the input data with a unified data structure, and the inference framework can convert the input data into a data structure supported by the inference model based on the actually deployed inference model; at the same time, after the inference framework outputs the output data of the seventh data structure, it can also be converted into the service data of the eighth data structure; thus, for the framework user, based on the unified upper-layer code, that is, using the unified data structure, the inference framework can convert the unified data structure into a model data structure supported by different inference models through the basic function module, improving the parallelism of software development.

[0026] In some embodiments, the basic function module is used to provide at least one of the following components: a cross-platform function component, a plug-in management component, and a log management component, where: the cross-platform function component is used to provide a function library for each of the multiple system functions; the function library includes the system functions corresponding to at least one system platform; the plug-in management component is used to manage the life cycle of the plug-ins in the plug-in definition module; the log management component is used to monitor the life cycle change events of the plug-ins in the plug-in definition module and generate corresponding logs based on the life cycle change events.

[0027] Based on the above embodiments, since function libraries of each of the multiple system functions are provided, after the inference framework is deployed to any system platform, the system functions corresponding to the any system platform can be provided through the cross-platform function component, thereby achieving adaptation to multiple system platforms. At the same time, through the plugin management component, corresponding lifecycles can be set for the function plugins corresponding to each inference device, thereby uniformly managing all the function plugins corresponding to each inference device. At the same time, through the log management component, operation events such as the system time of the lifecycle change event, the plugin identifier of the function plugin operated, and the lifecycle change situation can be recorded, which is convenient for problem tracing.

[0028] On the other hand, an embodiment of the present disclosure provides an inference method, which is applied to an inference framework that provides an external plugin interface. The method includes: receiving an inference request based on a target inference device; in response to the inference request, through the external plugin interface provided by the inference framework, calling the internal plugin interface corresponding to the target inference device among the internal plugin interfaces of multiple inference devices provided by the inference framework to convert the data to be inferred into service data.

[0029] In some embodiments, the external plugin interface includes an external device interface, an external processing interface, and an external inference interface. The step of, through the external plugin interface provided by the inference framework, calling the internal plugin interface corresponding to the target inference device among the internal plugin interfaces of multiple inference devices provided by the inference framework to convert the data to be inferred into service data includes: based on the external device interface provided by the inference framework, calling the internal device interface corresponding to the target inference device in the inference framework to load an inference model in the target inference device; based on the external processing interface provided by the inference framework, calling the internal processing interface corresponding to the target inference device in the inference framework to process the data to be inferred to obtain input data and input the input data into the inference model; based on the external inference interface provided by the inference framework, calling the internal inference interface corresponding to the target inference device in the inference framework to use the inference model to convert the input data into output data; based on the external processing interface provided by the inference framework, calling the internal processing interface corresponding to the target inference device in the inference framework to convert the output data into the service data; based on the external device interface provided by the inference framework, calling the internal device interface corresponding to the target inference device in the inference framework to release the inference model in the target inference device.

[0030] In another aspect, embodiments of the present disclosure provide an inference device, which is applied to an inference framework that provides an external plug-in interface. The inference device includes: a receiving module, configured to receive an inference request based on a target inference device; and an inference module, configured to, in response to the inference request, call the internal plug-in interface corresponding to the target inference device among the internal plug-in interfaces of a plurality of inference devices provided by the inference framework through the external plug-in interface provided by the inference framework to convert the data to be inferred into service data.

[0031] In yet another aspect, embodiments of the present disclosure provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0032] In yet another aspect, embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method.

[0033] In yet another aspect, embodiments of the present disclosure provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes to implement some or all of the steps in the above method.

[0034] In yet another aspect, embodiments of the present disclosure provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method.

[0035] Based on the above embodiments, since the inference framework encapsulates the inference processes of different inference devices into corresponding internal plug-in interfaces through a plug-in definition module, and defines a unified external plug-in interface corresponding to different inference devices through the basic function module, during the process of executing an inference task, an inference request can be received based on the external plug-in interface, and the internal plug-in interface corresponding to the target inference device can be called to implement the corresponding inference process, which can reduce the coupling degree between the inference functions of different inference devices and the inference framework and improve the stability of the inference framework. At the same time, since the inference functions of different inference devices are deployed in the plug-in definition module in the form of function plug-ins, in the case of needing to add a new inference device, only the function plug-in corresponding to the inference device needs to be defined, that is, the internal plug-in interface corresponding to the inference device needs to be defined, and the inference function of the inference device can be supported. And since it is the internal plug-in interface that is changed, for framework users, there is no need to change the upper-layer code of the framework, which not only reduces the development difficulty but also saves development costs and expands the application scope of the framework.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0038] Figure 1 Schematic diagram of a framework of an inference framework provided by an embodiment of the present disclosure;

[0039] Figure 2 Schematic diagram of a framework of an inference framework provided by an embodiment of the present disclosure;

[0040] Figure 3 Schematic diagram of a framework of an inference framework provided by an embodiment of the present disclosure;

[0041] Figure 4 Schematic diagram of a framework of an inference framework provided by an embodiment of the present disclosure;

[0042] Figure 5 Schematic diagram of a framework of an inference framework provided by an embodiment of the present disclosure;

[0043] Figure 6 Schematic diagram of the implementation process of an inference method provided by an embodiment of the present disclosure;

[0044] Figure 7 Schematic diagram of the implementation process of an inference method provided by an embodiment of the present disclosure;

[0045] Figure 8 Schematic diagram of a system architecture provided by an embodiment of the present disclosure;

[0046] Figure 9 Schematic diagram of the association relationship between a functional plug-in, a basic function module, and an interface provided by an embodiment of the present disclosure;

[0047] Figure 10 Schematic diagram of interface calls during an inference process provided by an embodiment of the present disclosure;

[0048] Figure 11 Schematic diagram of the composition structure of an inference device provided by an embodiment of the present disclosure;

[0049] Figure 12 Schematic diagram of the hardware entity of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0051] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs. The terms used herein are only for the purpose of describing the present disclosure and are not intended to limit the present disclosure.

[0053] Figure 1 A framework schematic diagram of an inference framework provided for an embodiment of the present disclosure is shown in Figure 1 As shown, the inference framework 100 includes a plug-in definition module 110 and a basic function module 120. Among them,

[0054] The plug-in definition module 110 includes function plug-ins corresponding to each of the multiple inference devices; the function plug-ins are used to encapsulate the inference process into an internal plug-in interface;

[0055] The basic function module 120 is used to nest and combine the internal plug-in interfaces corresponding to each inference device to generate corresponding external plug-in interfaces; the external plug-in interfaces are used to respond to an inference request based on a target inference device and call the internal plug-in interface corresponding to the target inference device to convert the data to be inferred into service data.

[0056] Among them, the inference process may include the following processes: initialization, pre-processing, performing inference, post-processing, and resource recovery, etc. The above function plug-ins can encapsulate the inference process into an internal plug-in interface.

[0057] In some embodiments, the inference process may be encapsulated into at least one internal plug-in interface according to different links in the inference process. When there is one internal plug-in interface, the functional plug-in may encapsulate the above-mentioned initialization, pre-processing, inference execution, post-processing, and resource recycling together to obtain one internal plug-in interface; when there are three internal plug-in interfaces, the functional plug-in may encapsulate the above-mentioned initialization process and resource recycling process into the first internal plug-in interface, encapsulate the above-mentioned pre-processing process and post-processing process into the second internal plug-in interface, and use the above-mentioned inference execution process as the third internal plug-in interface; when there are five internal plug-in interfaces, each of the above-mentioned processes may be encapsulated into one internal plug-in interface respectively. Currently, the inference process may also be encapsulated into two, four, or other numbers, which are not listed one by one in this disclosure.

[0058] In some embodiments, the plug-in definition module in the inference framework may include functional modules corresponding to each inference device among multiple inference devices, that is, the inference framework includes internal plug-in interfaces corresponding to each inference device. Among them, the multiple inference devices include inference devices corresponding to different inference engines.

[0059] In some embodiments, the plug-ins in the embodiments of this disclosure are dynamically inserted instead of statically written during programming. A plug-in is a type of functional component, and any functional component with a reserved interface in an application program is a plug-in.

[0060] It should be noted that since the inference process corresponding to each inference device is implemented through the functional plug-ins corresponding to each inference device in the embodiments of this disclosure, when it is necessary to add a new inference device to the inference architecture, the support of the inference architecture for the inference device can be realized by adding the functional plug-in corresponding to the inference device.

[0061] In some embodiments, the inference framework further includes a basic function module 120, and the basic function module 120 provides an external plug-in interface. Among them, there is a corresponding relationship between the external plug-in interface and the internal plug-in interfaces corresponding to each inference device. After the external plug-in interface receives an inference request based on a target inference device, it may call the internal plug-in interface corresponding to the target inference device from the internal plug-in interfaces corresponding to the multiple inference devices to process the inference request.

[0062] Exemplarily, the inference framework supports the inference processes of the first inference device, the second inference device, and the third inference device. That is, the plugin definition module in the inference framework includes a first function plugin corresponding to the first inference device, a second function plugin corresponding to the second inference device, and a third function plugin corresponding to the third inference device. Based on the function plugins corresponding to each of the above inference devices, the plugin definition module can provide a first internal plugin interface, a second internal plugin interface, and a third internal plugin interface. Correspondingly, the basic function module 120 can provide an external plugin interface corresponding to each of the first internal plugin interface, the second internal plugin interface, and the third internal plugin interface. After the external plugin interface receives an inference request based on the first inference device, it can call the first internal plugin interface to process the inference request; after the external plugin interface receives an inference request based on the second inference device, it can call the second internal plugin interface to process the inference request; after the external plugin interface receives an inference request based on the third inference device, it can call the third internal plugin interface to process the inference request. Among them, the embodiments of the present disclosure can also support other numbers of inference devices.

[0063] In some other embodiments, the function plugin corresponding to each inference device may include multiple function sub-plugins. That is, for each inference device, the inference process based on the inference device can be implemented through the multiple function sub-plugins corresponding to the inference device. Correspondingly, for each function sub-module, the basic function module can provide an external plugin interface corresponding to each function sub-module, and the external plugin interface can call the function sub-module of each inference device.

[0064] Exemplarily, when the inference framework supports the inference processes of the first inference device, the second inference device, and the third inference device, and the function plug-in includes the first function sub-plug-in, the second function sub-plug-in, and the third function sub-plug-in, the plug-in definition module can provide three types of internal plug-in interfaces, including the first type of internal plug-in interface corresponding to the first function sub-plug-in, the second type of internal plug-in interface corresponding to the second function sub-plug-in, and the third type of internal plug-in interface corresponding to the third function sub-plug-in. Among them, the first type of internal plug-in interface includes the first internal plug-in interface corresponding to each inference device, the second type of internal plug-in interface includes the second internal plug-in interface corresponding to each inference device, and the third type of internal plug-in interface includes the third internal plug-in interface corresponding to each inference device. Correspondingly, the basic function module 120 can provide the first external plug-in interface corresponding to the first type of internal plug-in interface, the second external plug-in interface corresponding to the second type of internal plug-in interface, and the third external plug-in interface corresponding to the third type of internal plug-in interface. After receiving an inference request based on the first inference device, the inference request may include a first sub-request, a second sub-request, and a third sub-request. Among them, the first sub-request can be received through the first external plug-in interface, and then the first internal plug-in interface corresponding to the first inference device can be called to process the first sub-request; the second sub-request can be received through the second external plug-in interface, and then the second internal plug-in interface corresponding to the first inference device can be called to process the second sub-request; the third sub-request can be received through the third external plug-in interface, and then the third internal plug-in interface corresponding to the first inference device can be called to process the third sub-request. Among them, the embodiments of the present disclosure can also support other numbers of inference devices, and the function plug-in can include other numbers of function sub-plug-ins.

[0065] Based on the above embodiments, since the inference framework encapsulates the inference processes of different inference devices into corresponding internal plug-in interfaces through the plug-in definition module, and, through the basic function module, defines unified external plug-in interfaces corresponding to different inference devices. During the execution of the inference task, an inference request can be received based on the external plug-in interface, and the internal plug-in interface corresponding to the target inference device can be called to implement the corresponding inference process, which can reduce the coupling degree between the inference functions of different inference devices and the inference framework, and improve the stability of the inference framework; at the same time, since the inference functions of different inference devices are deployed in the plug-in definition module in the form of function plug-ins, in the case of needing to add a new inference device, only the function plug-in corresponding to the inference device needs to be defined, that is, the internal plug-in interface corresponding to the inference device is defined, and the inference function of the inference device can be supported; and since what is changed is the internal plug-in interface, for the framework user, there is no need to change the upper-layer code of the framework, which can reduce the development difficulty and save the development cost while improving the application scope of the framework.

[0066] Figure 2 A framework schematic diagram of an inference framework provided by an embodiment of the present disclosure, as Figure 2 shown, the inference framework 100 includes a plugin definition module 110 and a basic function module 120. Among them, the function plugin includes a device plugin 111 for providing an internal device interface. The internal plugin interface includes an internal device interface, and the external plugin interface includes an external device interface. Among them, the basic function module nests and combines the internal device interfaces to generate an external device interface; when the inference request includes a device management request for the target inference device, the external device interface is used to call the internal device interface corresponding to the target inference device based on the device management request to manage the target inference device.

[0067] Based on the above embodiment, since the basic function module provides an external device interface externally, and this external device interface can call the internal device interfaces of multiple inference devices provided by the device plugin, thus, hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external device interface is provided externally, thereby providing support for large-scale production in the software industry, shortening the R & D cycle and saving R & D costs; at the same time, since the internal device interfaces of each inference device are provided by the device plugin included in the plugin definition module, when a new device is connected, only the device plugin corresponding to the inference device needs to be defined, that is, the internal device interface corresponding to the inference device is defined, and the management of the new device can be supported. And since it is the internal device interface that changes, for the framework user, there is no need to change the upper-layer code of the framework, further saving development costs and expanding the application scope of the framework.

[0068] In some embodiments, the internal device interface includes a device binding sub-interface, a device status acquisition sub-interface, and a device memory operation sub-interface; where: the device binding sub-interface is used to obtain the identification information of the inference device; the device status acquisition sub-interface is used to obtain the device status of the inference device; the device memory operation sub-interface is used to operate the device memory.

[0069] Based on the above embodiment, since the device plugin can provide sub-interfaces with different functions for the inference device, during the process of defining the interfaces for device management for the inference device, the various functions of the inference device management can be separated, effectively reducing the coupling degree between different function sub-interfaces and ensuring the stability of device management during the inference process; at the same time, during the process of using the inference device to process inference tasks, if an inference error occurs, the problem location can be quickly located, and the developer only needs to redefine the sub-interface where the current problem appears, reducing the development cost.

[0070] Figure 3A framework schematic diagram of an inference framework provided by an embodiment of the present disclosure is as follows Figure 3 As shown, the inference framework 100 includes a plug-in definition module 110 and a basic function module 120. Among them, the function plug-in includes a processing plug-in 112 for providing an internal processing interface. The internal plug-in interface includes an internal processing interface, and the external plug-in interface includes an external processing interface. Among them, the basic function module nests and combines the internal processing interfaces to generate an external processing interface. When the inference request includes a pre-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the pre-processing request, and convert the data to be inferred into the input data of the inference model. When the inference request includes a post-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the post-processing request, and convert the output data of the inference model into service data.

[0071] Based on the above embodiment, since the basic function module provides an external processing interface externally, this external processing interface can call the internal processing interfaces of multiple inference devices provided by the processing plug-in. Therefore, hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external processing interface is provided externally, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal processing interfaces of each inference device are provided by the processing plug-in included in the plug-in definition module, when a new device is connected, only the processing plug-in corresponding to the inference device needs to be defined, that is, the internal processing interface corresponding to the inference device is defined, and the data processing of the inference process can be completed using the inference device, improving the scalability and deployment efficiency of the inference framework.

[0072] In some embodiments, the internal processing interface includes a pre-processing sub-interface and a post-processing sub-interface. The pre-processing sub-interface is used to convert the data to be inferred into the input data of the inference model. The post-processing sub-interface is used to convert the output data of the inference model into service data.

[0073] Based on the above embodiment, since the processing plug-in can provide sub-interfaces with different functions for the inference device, the pre-processing process and the post-processing process in the inference process can be separated during the interface definition process for the processing of the inference device, effectively reducing the coupling degree between different function sub-interfaces and ensuring the stability of the data processing process during the inference process. At the same time, during the process of using the inference device to process the inference task, if an inference error occurs, the problem location can be quickly located. During the process of eliminating the problem, developers only need to re-define the sub-interface where the current problem occurs, reducing the development cost.

[0074] Figure 4A framework schematic diagram of an inference framework provided by an embodiment of the present disclosure is as follows Figure 4 As shown, the inference framework 100 includes a plug-in definition module 110 and a basic function module 120. Among them, the function plug-in includes an inference plug-in 113 for providing an internal inference interface. The internal plug-in interface includes an internal inference interface, and the external plug-in interface includes an external inference interface. Among them, the basic function module nests and combines the internal inference interfaces to generate an external inference interface. When the inference request includes a model inference request for the target inference device, the external inference interface is used to call the corresponding internal inference interface of the target inference device based on the model inference request, and convert the input data into output data based on the inference model.

[0075] Based on the above embodiment, since the external inference interface is provided externally through the basic function module, this external inference interface can call the internal inference interfaces of multiple inference devices provided by the inference plug-in. Therefore, hardware differences can be isolated at the framework level. While supporting multiple inference devices, a unified external inference interface is provided externally, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal inference interfaces of each inference device are provided by the inference plug-ins included in the plug-in definition module, when a new device is connected, only the inference plug-in corresponding to the inference device needs to be defined, that is, the internal inference interface corresponding to the inference device is defined, and the inference process can be executed using this inference device, which improves the scalability and deployment efficiency of the inference framework.

[0076] In some embodiments, the internal inference interface includes internal inference interfaces corresponding to at least one inference model. The external inference interface is used to call the corresponding target internal inference interface of the target inference device based on the model inference request for the target inference device, and convert the input data into output data based on the target inference model. The inference request is used to call the target inference model in the at least one inference model, and the target inference interface is used to manage the inference process of the target inference model.

[0077] Based on the above embodiments, since an external processing interface is provided externally through the basic function module, and this external processing interface can call the internal processing interfaces corresponding to at least one inference model provided by the processing plug-in, thus, model differences can be isolated at the framework level. While supporting multiple inference devices and inference models, a unified external inference interface is provided externally, thereby shortening the R & D cycle and saving R & D costs. At the same time, since the internal inference interfaces corresponding to each inference model are provided by the inference plug-ins included in the plug-in definition module, in the case of needing to adapt to a new inference model, only the inference plug-in corresponding to the new inference model needs to be defined, that is, the internal processing interface corresponding to the new inference model is defined, and then the input data can be converted into output data based on the new inference model, improving the scalability and deployment efficiency of the inference framework.

[0078] In some embodiments, the internal inference interface includes: a model parsing sub-interface, a model configuration sub-interface, a model inference sub-interface, and a result obtaining sub-interface. Among them, the model parsing sub-interface is used to obtain a model file and perform parsing. The model configuration sub-interface is used to configure the parsed model to obtain a configured model. The model inference sub-interface is used to complete the model inference process based on the configured model. The result obtaining sub-interface is used to obtain the inference result of the model inference process.

[0079] Based on the above embodiments, since the inference plug-in can provide sub-interfaces with different functions for the inference device, during the process of defining the interfaces for the inference process for the inference device, the processes of model parsing, model configuration, model inference, and result obtaining in the inference process can be separated, effectively reducing the coupling degree between different functional sub-interfaces and ensuring the stability of the inference process based on the inference model. At the same time, during the process of using the inference device to process the inference task, if an inference error occurs, the problem location can be quickly located. During the process of eliminating the problem, the developer only needs to redefine the sub-interface where the current problem occurs, reducing the development cost.

[0080] In some embodiments, based on the inference framework described in any of the above embodiments, the basic function module is used to provide at least one business function. When the business function is called, at least one of the following is executed: when the business function is called by the internal inference interface, converting the input data into the output data; when the business function is called by the internal processing interface, converting the data to be inferred into the input data of the inference model; when the business function is called by the internal processing interface, converting the output data of the inference model into business data.

[0081] Based on the above embodiments, since the service functions provided by the functional functions enable each internal functional plug-in to reuse the service functions, while improving the code reuse rate, the development difficulty can also be reduced.

[0082] In some embodiments, based on the inference framework described in any of the above embodiments, the basic functional module is used to provide an exchange data structure corresponding to each of the inference devices; the inference request is used to convert the data to be inferred in the first data structure into service data in the second data structure; when the internal processing interface is called, it is further used to convert the data to be inferred in the first data structure into data to be inferred in the third data structure; when the internal processing interface is called, it is further used to convert the service data in the fourth data structure into the service data in the second data structure; the third data structure and the fourth data structure are data structures supported by the target inference device.

[0083] Based on the above embodiments, since the basic functional module provides an exchange data interface corresponding to each inference device, in the process of using different inference devices for data processing, the framework user can input the data to be inferred (the first data structure) with a unified data structure, and the inference framework can convert the data to be inferred into a data structure supported by the inference device (the third data structure) based on the actually deployed inference device; at the same time, after the inference framework outputs the service data in the fourth data structure, it can also be converted into the service data in the second data structure; thus, for the framework user, based on the unified upper-layer code, that is, using the unified data structure, the inference framework can convert the unified data structure into an exchange data structure supported by different inference devices through the basic functional module, improving the parallelism of software development.

[0084] In some embodiments, based on the inference framework described in any of the above embodiments, the basic functional module is used to provide a model data structure corresponding to each of the inference models; when the internal inference interface is called, it is further used to convert the input data in the fifth data structure into input data in the sixth data structure; the internal inference interface is used to convert the input data in the sixth data structure into output data in the seventh data structure based on the target inference model; when the internal inference interface is called, it is further used to convert the output data in the seventh data structure into the service data in the eighth data structure.

[0085] Based on the above embodiments, since the basic function module provides the model data structure corresponding to each of the inference models, during the process of performing model inference using different inference models, the framework user can input input data with a unified data structure, and the inference framework can convert the input data into the data structure supported by the inference model based on the actually deployed inference model; meanwhile, after the inference framework outputs the output data of the seventh data structure, it can also be converted into the service data of the eighth data structure; thus, for the framework user, based on the unified upper-layer code, that is, using the unified data structure, the inference framework can convert the unified data structure into the model data structures supported by different inference models through the basic function module, improving the parallelism of software development.

[0086] Figure 5 The framework schematic diagram of an inference framework provided by an embodiment of the present disclosure is as follows Figure 5 As shown, the inference framework 100 includes a plug-in definition module 110 and a basic function module 120. Among them, the basic function module 120 includes at least one of the following components: a cross-platform function component 121, a plug-in management component 122, and a log management component 123, where:

[0087] The cross-platform function component 121 is used to provide the function library of each of the multiple system functions; the function library includes the system functions corresponding to at least one system platform.

[0088] In some embodiments, the system functions may include general functions such as read functions and write functions. For each system function, the cross-platform function component can provide the system function corresponding to each system platform for the system function. After the inference framework is deployed to any system platform, the system functions corresponding to the any system platform can be provided through the cross-platform function component, thereby realizing the adaptation to multiple system platforms.

[0089] The plug-in management component 122 is used to manage the life cycle of the plug-ins of the plug-in definition module.

[0090] In some embodiments, the plug-in management component can manage the life cycle of the plug-ins in the plug-in definition module based on the dimension of the inference device, that is, the plug-in management component can set the corresponding life cycle for the function plug-ins corresponding to each inference device, and then uniformly manage all the function plug-ins corresponding to each inference device.

[0091] Exemplarily, the inference framework includes function plugins corresponding to a first inference device and plugins corresponding to a second inference device. The function plugins may include a first function sub-plugin and a second function sub-plugin. Based on the above embodiments, the plugin management component may manage the lifecycle of the plugins in the plugin definition module based on the inference device, that is, the plugin management component may set a unified lifecycle for the first function sub-plugin and the second function sub-plugin corresponding to the first inference device, and set a unified lifecycle for the first function sub-plugin and the second function sub-plugin corresponding to the second inference device, thereby implementing the lifecycle management in terms of the inference device dimension.

[0092] In some embodiments, the plugin management component may manage the lifecycle of the plugins in the plugin definition module based on the function plugin dimension, that is, the plugin management component may set a corresponding lifecycle for each type of function plugin, and thereby uniformly manage the function plugins of each inference device corresponding to the type.

[0093] Exemplarily, the inference framework includes function plugins corresponding to a first inference device and plugins corresponding to a second inference device. The function plugins may include a first function sub-plugin and a second function sub-plugin. Based on the above embodiments, the plugin management component may manage the lifecycle of the plugins in the plugin definition module based on the function sub-plugin, that is, the plugin management component may set a unified lifecycle for the first function sub-plugin corresponding to the first inference device and the first function sub-plugin corresponding to the second inference device, and set a unified lifecycle for the second function sub-plugin corresponding to the first inference device and the second function sub-plugin corresponding to the second inference device, thereby implementing the lifecycle management in terms of the function plugin dimension.

[0094] In some embodiments, the plugin management component may simultaneously manage the lifecycle of the plugins in the plugin definition module based on both the inference device dimension and the function plugin dimension, that is, the plugin management component may set a corresponding lifecycle for each function sub-plugin, and thereby precisely manage the lifecycle of different function sub-plugins in different inference devices.

[0095] Exemplarily, the inference framework includes function plugins corresponding to a first inference device and plugins corresponding to a second inference device. The function plugins may include a first function sub-plugin and a second function sub-plugin. Based on the above embodiments, the plugin management component may set a corresponding lifecycle for the first function sub-plugin corresponding to the first inference device; set a corresponding lifecycle for the first function sub-plugin corresponding to the second inference device; set a corresponding lifecycle for the second function sub-plugin corresponding to the first inference device; set a corresponding lifecycle for the second function sub-plugin corresponding to the second inference device, thereby precisely managing the lifecycle of different function sub-plugins in different inference devices.

[0096] The log management component 123 is used to monitor the life cycle change events of the plugins in the plugin definition module and generate corresponding logs based on the life cycle change events.

[0097] In some embodiments, the log management component 123 can be used to monitor the life cycle change events corresponding to each functional plugin in the above plugin definition module. The life cycle change events are generated in response to life cycle management requests for managing the life cycle of each functional plugin, including but not limited to extension operations, shortening operations, etc. In response to the life cycle change event, the log management component can record the system time of the life cycle change event, the plugin identifier of the functional plugin operated, the life cycle change situation, etc.

[0098] Based on the above embodiments, since function libraries of each of the multiple system functions are provided, after the inference framework is deployed to any system platform, the cross-platform function component can provide the system functions corresponding to the any system platform, thereby achieving adaptation to multiple system platforms; meanwhile, through the plugin management component, corresponding life cycles can be set for the functional plugins corresponding to each inference device, thereby uniformly managing all the functional plugins corresponding to each inference device; meanwhile, through the log management component, operation events such as the system time of the life cycle change event, the plugin identifier of the functional plugin operated, and the life cycle change situation can be recorded, facilitating problem tracing.

[0099] Embodiments of the present disclosure provide an inference method, which can be executed by a processor of a computer device. Herein, the computer device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable gaming device), etc.

[0100] Figure 6 For an implementation flow diagram of an inference method provided by embodiments of the present disclosure, as Figure 6 shown, the method includes the following steps S601 to step S602:

[0101] Step S601, receiving an inference request based on a target inference device.

[0102] Step S602, in response to the inference request, through an external plugin interface provided by the inference framework, calling an internal plugin interface corresponding to the target inference device among internal plugin interfaces of multiple inference devices provided by the inference framework to convert the data to be inferred into service data.

[0103] In some embodiments, the inference framework may include internal plugin interfaces corresponding to each inference device among a plurality of inference devices. Among them, the internal plugin interface is an interface obtained by encapsulating the inference process with a functional plugin included in the plugin definition module in the inference framework, and the external plugin interface is generated by the basic functional module in the inference framework by nesting and combining the internal plugin interfaces. Among them, the inference process may include the following processes: initialization, pre-processing, performing inference, post-processing, and resource recycling. The above functional plugin may encapsulate the inference process into an internal plugin interface.

[0104] In some embodiments, the inference process may be encapsulated into at least one internal plugin interface according to different processes in the inference process. When there is one internal plugin interface, the functional plugin may encapsulate the above initialization, pre-processing, performing inference, post-processing, and resource recycling together to obtain one internal plugin interface; when there are three internal plugin interfaces, the functional plugin may encapsulate the above initialization process and resource recycling process into the first internal plugin interface, encapsulate the above pre-processing process and post-processing process into the second internal plugin interface, and use the above performing inference process as the third internal plugin interface; when there are five internal plugin interfaces, each of the above processes may be encapsulated into one internal plugin interface respectively. Currently, the inference process may also be encapsulated into two, four, or other numbers, which are not listed one by one in this disclosure.

[0105] In some embodiments, the plugin definition module in the inference framework may include functional modules corresponding to each inference device among a plurality of inference devices, that is, the inference framework includes internal plugin interfaces corresponding to each inference device. Among them, the plurality of inference devices include inference devices corresponding to different inference engines.

[0106] In some embodiments, the plugins in the embodiments of the present disclosure are dynamically inserted instead of being statically written during programming. A plugin is a type of functional component, and any functional component with an interface reserved in the application program is a plugin.

[0107] In some embodiments, the external plugin interface includes at least one of the following: an external device interface, an external processing interface, and an external inference interface.

[0108] When the external plug-in interface includes an external device interface, the external device interface is generated by nesting and combining the internal device interfaces of the above basic function modules, and the internal device interfaces are the interfaces provided by the device plug-ins included in the above plug-in definition module. That is, the plug-in definition module in the inference framework can provide the internal device interface corresponding to each inference device, and at the same time, can provide an external device interface through the basic function module in the inference framework. When the external device interface receives a device management request for the target inference device, it can determine the internal device interface corresponding to the target inference device among the internal device interfaces corresponding to multiple inference devices, and process the device management request for the target inference device based on the internal device interface corresponding to the target inference device.

[0109] In some embodiments, the internal device interface includes at least one of the following sub-interfaces: a device binding sub-interface, a device status acquisition sub-interface, and a device memory operation sub-interface. Among them, when the device management request is used to obtain the identification information of the inference device, the device binding sub-interface of the inference framework can be called to obtain the identification information of the inference device; when the device management request is used to obtain the device status of the inference device, the device status acquisition sub-interface of the inference framework can be called to obtain the device status of the inference device; when the device management request is used to manage the memory of the inference device, the device memory operation sub-interface of the inference framework can be called to operate the memory of the inference device.

[0110] Exemplarily, during the initialization process in the inference process, the device identification of the inference device to be initialized can be determined based on the internal device interface, and the inference device to be initialized can be determined among multiple inference devices based on the above device binding sub-interface. At the same time, the inference model can be loaded into the memory of the inference device based on the above device memory operation sub-interface. In some embodiments, the device status of the inference device can also be obtained based on the above device status acquisition sub-interface. When the device status indicates that the inference device cannot currently process the inference request, corresponding prompt information can be fed back.

[0111] Exemplarily, during the resource release process in the inference process, the device identification of the inference device for which resources need to be released can be determined based on the internal device interface, and the inference device for which resources need to be released can be determined among multiple inference devices based on the above device binding sub-interface. At the same time, the memory loaded with the inference model can be released based on the above device memory operation sub-interface. In some embodiments, the device status of the inference device can also be obtained based on the above device status acquisition sub-interface. When the device status indicates that the inference device is currently executing the inference process, the current inference progress can be fed back.

[0112] When the external plug-in interface includes an external inference interface, the external inference interface is generated by the nested combination of the above basic function modules for the internal inference interface, and the internal inference interface is the interface provided by the inference plug-in included in the above plug-in definition module. That is, the plug-in definition module in the inference framework can provide the internal inference interface corresponding to each inference device, and at the same time, can provide an external inference interface through the basic function module in the inference framework. The external inference interface can, when receiving a model inference request for the target inference device, determine the internal inference interface corresponding to the target inference device among the internal inference interfaces corresponding to multiple inference devices, and process the model inference request for the target inference device based on the internal inference interface corresponding to the target inference device.

[0113] In some embodiments, the internal inference interface includes the internal inference interfaces corresponding to at least one inference model. The above model inference request for the target inference device may include a model inference request for a target inference model among the at least one inference models. Correspondingly, after the model inference request for the target inference model based on the external inference interface, the internal inference interface corresponding to the target inference model can be determined among the internal inference interfaces corresponding to multiple inference models, and the internal inference interface corresponding to the target inference model can be called to process the model inference request for the target inference device.

[0114] In some embodiments, the internal device interface includes at least one of the following sub-interfaces: a parsing model sub-interface, a model configuration sub-interface, a model inference sub-interface, and a result acquisition sub-interface. Among them, when the model inference request is used to parse the inference model, the parsing model sub-interface of the inference framework can be called to obtain the model file of the inference model and perform parsing; when the model inference request is used to configure the inference model, the model configuration sub-interface of the inference framework can be called to configure the parsed model to obtain a configured model; when the device management request is used to execute the inference process, the model inference sub-interface of the inference framework can be called to complete the model inference process based on the configured model; when the device management request is used to obtain the inference result, the result acquisition sub-interface of the inference framework can be called to obtain the inference result of the model inference process.

[0115] Exemplarily, during the execution of the inference process, the model file of the target inference model can be obtained and parsed based on the parsing model sub-interface, and then the parsed model can be configured through the model configuration sub-interface to obtain a configured model. The configured model is controlled to complete the model inference process through the model inference sub-interface, and finally the inference result of the inference process is obtained through the result acquisition sub-interface.

[0116] Figure 7It is an optional process schematic diagram of the inference method provided by an embodiment of the present disclosure, and this method can be executed by a processor of a computer device. The external plug-in interface includes an external device interface, an external processing interface, and an external inference interface; based on Figure 6 , Figure 6 S602 in can include S701 to S703, which will be described in combination with the steps shown in Figure 7 .

[0117] Step S701: Based on the external device interface provided by the inference framework, call the internal device interface corresponding to the target inference device in the inference framework, and load the inference model in the target inference device;

[0118] Step S702: Based on the external processing interface provided by the inference framework, call the internal processing interface corresponding to the target inference device in the inference framework, process the data to be inferred to obtain input data, and input the input data into the inference model;

[0119] Step S703: Based on the external inference interface provided by the inference framework, call the internal inference interface corresponding to the target inference device in the inference framework, and use the inference model to convert the input data into output data;

[0120] Step S704: Based on the external processing interface provided by the inference framework, call the internal processing interface corresponding to the target inference device in the inference framework, and convert the output data into the service data;

[0121] Step S705: Based on the external device interface provided by the inference framework, call the internal device interface corresponding to the target inference device in the inference framework, and release the inference model in the target inference device.

[0122] Next, the application of the inference method provided by the embodiment of the present disclosure in an actual scenario will be described. Taking a system for collaboratively developing a scalable architecture for quickly implementing model inference functions as an example for description.

[0123] In the existing inference framework, all device management and model inference logics are together, and different interfaces are provided for each function externally, which is not easy to replace and maintain. This solution decouples device hardware and model inference through a plug-in system. The caller does not need to understand the internal interfaces and can use a unified interface to call different functions, reducing the design and development difficulty and improving the parallelism of software development.

[0124] Please refer to Figure 8 , which shows a schematic diagram of a system architecture. As shown in Figure 8As shown in the figure, above the inference framework layer 800, there may be a plug-in definition module 810 and a basic function module 820. Among them, the plug-in definition module 810 can be used to provide three types of interfaces, including a device management plug-in interface 811 (corresponding to the internal device interface in the above embodiment), a model inference plug-in interface 812 (corresponding to the internal inference interface in the above embodiment), and a general function plug-in interface 813 (corresponding to the internal processing interface in the above embodiment); the basic function module 820 can include a plug-in management component 821, a cross-platform library component 822, and a log management component 823.

[0125] In some embodiments, the basic function module 820 can be used to provide: common functions in the business dimension (corresponding to at least one business function in the above embodiment), general data structures and general data exchange formats in the artificial intelligence industry (corresponding to the exchange data structure corresponding to each of the above-mentioned inference devices). Furthermore, it can decouple developers from the development platform and achieve efficient data flow by customizing the data exchange format and operations.

[0126] Among them, the common functions in the business dimension are common functions that can be shared among different plug-ins, and these common functions are related to the business category. Exemplarily, if there are plug-in 1, plug-in 2, and plug-in 3, and all three of these plug-ins are for image processing, at this time, the common functions in the business dimension can be functions such as scaling and color conversion. These common functions can be placed in the basic function module, and plug-in 1, plug-in 2, and plug-in 3 can all use these functions. Correspondingly, each plug-in does not need to redefine these common functions. Another example is that if plug-in 1, plug-in 2, and plug-in 3 are all for linear processing, at this time, the common functions in the business dimension can be functions related to linear algebra, and then the functions related to linear algebra can be placed in the basic function module.

[0127] Among them, the general data exchange format is the input data structure and output data structure supported by different inference devices. Exemplarily, before inputting the data of the original data structure into the inference device, the data of the original data structure can be converted into the input data structure supported by the inference device, and then the data can be input into the inference device; correspondingly, after the inference device obtains the data of the output data structure and before outputting the data, the data of the output data structure can be converted into the data of the target data structure required by the external interface. Based on this, for the inference processes of different inference devices, from the perspective of the inference framework user, the input data is all the data of the original data structure. The basic functional modules in the inference framework can provide the input data interfaces supported by each inference device, and then based on the target inference device used, the data of the original data structure can be converted into the data of the input data structure supported by the target inference device; after the target inference device obtains the data of the output data structure, the data of different output data structures obtained by different inference devices can also be converted into the data of a unified data structure and output. From the perspective of the inference framework user, without paying attention to the data structures supported by each inference device, the inference process based on different inference devices can be completed.

[0128] In some embodiments, the device management plug-in interface 811 is a general interface for device management and memory operations, thereby achieving the rapid access of different inference devices.

[0129] Among them, the device management plug-in interface 811 may include a device binding sub-interface, a device status acquisition sub-interface, and a device memory operation sub-interface; where: the device binding sub-interface is used to obtain the identification information of the inference device; the device status acquisition sub-interface is used to obtain the device status of the inference device; the device memory operation sub-interface is used to operate the device memory.

[0130] In some embodiments, the model inference plug-in interface 812 is a model inference interface for different inference devices, which is used to achieve the rapid access of different inference devices.

[0131] Among them, the model inference plug-in interface 812 may include a model parsing sub-interface, a model configuration sub-interface, a model inference sub-interface, and a result acquisition sub-interface; where, the model parsing sub-interface is used to obtain the model file and perform parsing; the model configuration sub-interface is used to configure the parsed model to obtain a configured model; the model inference sub-interface is used to complete the model inference process based on the configured model; the result acquisition sub-interface is used to obtain the inference result of the model inference process.

[0132] In some embodiments, the general function plug-in interface 813 is a general interface for pre-processing and post-processing of the model.

[0133] Among them, the general function plug-in interface 813 may include a pre-processing sub-interface and a post-processing sub-interface; the pre-processing sub-interface is used to convert the data to be inferred into the input data of the inference model; the post-processing sub-interface is used to convert the output data of the inference model into service data.

[0134] Please refer to Figure 9 , Figure 9 which shows a schematic diagram of the association relationship between a function plug-in, a basic function module, and an interface. It can be seen that the inference framework provided by the embodiments of the present disclosure may include three function plug-ins: a model inference plug-in 910, a general function plug-in 920, and a device management plug-in 930. Correspondingly, each function plug-in is used to provide a corresponding internal interface. For example, the model inference plug-in 910 provides a model inference interface 911, the general function plug-in 920 provides a general function interface 921, and the device management plug-in 930 provides a device management interface 931. The inference framework may further include a basic function module 940, and the basic function module 940 may call the model inference interface 911 to provide an external model inference interface (corresponding to the external inference interface in the above embodiments) to implement the model inference function; the basic function module 940 may call the general function interface 921 to provide an external general function interface (corresponding to the external processing interface in the above embodiments) to implement the data processing function; the basic function module 940 may call the device management interface 931 to provide an external device management interface (corresponding to the external device interface in the above embodiments) to implement the device management function.

[0135] In some embodiments, since the basic function module 940 can provide common functions, artificial intelligence industry general data structures, and general data exchange formats in the service dimension, the model inference plug-in 910 and the general function plug-in 920 can use the functions and / or data structures provided by the basic function module 940 in the process of implementing the corresponding functions.

[0136] In some embodiments, the inference process may include processes such as initialization, pre-processing, inference execution, post-processing, and resource recovery. Please refer to Figure 10 which shows a schematic diagram of interface calls during the inference process.

[0137] Among them, the above inference process may include:

[0138] Step S1001, the initialization process.

[0139] Step S1002, the pre-processing process.

[0140] Step S1003, the inference execution process.

[0141] Step S1004, the post-processing process.

[0142] Step S1005, the process of recycling resources.

[0143] Correspondingly, the above basic functional module can provide an external interface of the framework (corresponding to the external plug-in interface in the above embodiments), such as Figure 10 the external interface 1010 of the framework. The external interface 1010 of the framework may include an external device interface 1011, an external inference interface 1012, and an external processing interface 1013. In some embodiments, the external interface 1010 of the framework may also include basic data functions and functional modules for providing common functions in the above business dimension, general data structures in the artificial intelligence industry, and general data exchange formats.

[0144] In some embodiments, the initialization process in step S1001 may be implemented by calling the external device interface 1011; the pre-processing process in step S1002 may be implemented by calling the external processing interface 1013; the execution inference process in step S1003 may be implemented by calling the external inference interface 1012; the post-processing process in step S1004 may be implemented by calling the external processing interface 1013; the process of recycling resources in step S1005 may be implemented by calling the external device interface 1011.

[0145] Among them, the plug-in definition module in the inference framework can provide a device management plug-in 1021, a general function plug-in 1022, and a model inference plug-in 1023. Correspondingly, the device management plug-in 1021 can provide internal device interfaces 1031 corresponding to multiple inference devices, the general function plug-in 1023 can provide internal processing interfaces 1033 corresponding to multiple inference devices, and the model inference plug-in 1022 can provide internal inference interfaces 1032 corresponding to multiple inference devices. As Figure 10 shown, the basic functional module can generate the external device interface 1011 by nested combination of the internal device interfaces 1031 of multiple inference devices, can also generate the external processing interface 1013 by nested combination of the internal processing interfaces 1033 of multiple inference devices, and can also generate the external inference interface 1013 by nested combination of the internal inference interfaces 1032 of multiple inference devices.

[0146] Based on the above embodiments, the inference device and the inference model can be abstracted, the differences between the platform and the hardware can be isolated, a unified data structure for exchange can be provided, the software reuse rate and the parallelism of software development can be improved, support for large-scale production in the software industry can be provided. At the same time, the R & D cycle of the software can be shortened, the R & D cost can be saved, more flexibility can be brought to the program developers, and new plug-ins can be added and the existing functions can be improved after the software is released.

[0147] In some implementation scenarios, different developers can collaborate on development, which improves development efficiency. For example, device management developers can implement the device management functions of the corresponding platform based on the defined device management interfaces without the need to understand the technologies and businesses related to inference; model inference developers can implement the inference functions of the corresponding platform based on the defined inference interfaces; general function developers can implement different business functions based on the general function interfaces and basic functions without the need to concern themselves with platform and hardware differences; framework users can use the above-mentioned unified and stable interfaces to implement different new functions through different configurations. In the case of new hardware adaptation, only device management developers and model inference developers need to write the corresponding code; in the case of new model network interfaces, only general function developers need to adapt; in the case of new hardware or updated model network structures, only framework users need to modify the corresponding configurations.

[0148] Based on the foregoing embodiments, an embodiment of the present disclosure provides an inference device. The device includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. During implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0149] Figure 11 It is a schematic diagram of the composition structure of an inference device provided by an embodiment of the present disclosure. The inference device is applied to an inference framework, and the inference framework provides an external plugin interface, such as Figure 11 As shown, the inference device 1100 includes: a receiving module 1110 and an inference module 1120, where:

[0150] The receiving module 1110 is configured to receive an inference request based on a target inference device;

[0151] The inference module 1120 is configured to, in response to the inference request, call the internal plugin interface corresponding to the target inference device among the internal plugin interfaces of multiple inference devices provided by the inference framework through the external plugin interface provided by the inference framework to convert the data to be inferred into service data.

[0152] In some embodiments, the inference module 1120 is further configured to: based on the external device interface provided by the inference framework, call the internal device interface corresponding to the target inference device in the inference framework to load an inference model in the target inference device; based on the external processing interface provided by the inference framework, call the internal processing interface corresponding to the target inference device in the inference framework to process the data to be inferred to obtain input data, and input the input data into the inference model; based on the external inference interface provided by the inference framework, call the internal inference interface corresponding to the target inference device in the inference framework, and use the inference model to convert the input data into output data; based on the external processing interface provided by the inference framework, call the internal processing interface corresponding to the target inference device in the inference framework to convert the output data into the service data; based on the external device interface provided by the inference framework, call the internal device interface corresponding to the target inference device in the inference framework to release the inference model in the target inference device.

[0153] The description of the above device embodiments is similar to that of the above method embodiments, and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0154] It should be noted that in the embodiments of the present disclosure, if the above-mentioned inference method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0155] The embodiments of the present disclosure provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0156] Embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0157] Embodiments of the present disclosure provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes to implement some or all of the steps in the above method.

[0158] Embodiments of the present disclosure provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0159] It should be noted here that: the above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0160] Figure 12 The following is a schematic diagram of the hardware entity of an inference device provided by an embodiment of the present disclosure. As Figure 12 shown, the hardware entity of the inference device 1200 includes: a processor 1201 and a memory 1202. Among them, the memory 1202 stores a computer program that can run on the processor 1201. When the processor 1201 executes the program, the steps in the method of any of the above embodiments are implemented.

[0161] The memory 1202 stores a computer program that can run on a processor. The memory 1202 is configured to store instructions and applications executable by the processor 1201, and can also cache data to be processed or already processed by each module in the processor 1201 and the inference device 1200 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by a flash memory (FLASH) or a random access memory (Random Access Memory, RAM).

[0162] When the processor 1201 executes the program, it implements the steps of the inference method in any of the above. The processor 1201 generally controls the overall operation of the inference device 1200.

[0163] The embodiments of the present disclosure provide a computer storage medium that stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the inference method in any of the above embodiments.

[0164] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0165] The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices implementing the functions of the above processor are also possible, and the embodiments of the present disclosure do not make specific limitations.

[0166] The above computer storage medium / memory may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it may be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0167] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the sequence numbers of the above steps / processes do not indicate the order of execution, and the order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The sequence numbers of the embodiments of the present disclosure are only for description and do not represent the advantages or disadvantages of the embodiments.

[0168] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0169] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.

[0170] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] In addition, each functional unit in the embodiments of the present disclosure can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0172] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0173] Alternatively, if the above-mentioned integrated units of the present disclosure are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present disclosure. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.

[0174] As described above, it is only the implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure.

Claims

1. An inference framework system, characterized in that, the inference framework system includes a plug-in definition module and a basic function module; wherein, the plug-in definition module includes function plug-ins corresponding to each of the multiple inference devices; the function plug-ins are used to encapsulate the inference process into an internal plug-in interface; the basic function module is used to nestedly combine the internal plug-in interfaces corresponding to each inference device to generate corresponding external plug-in interfaces; the external plug-in interfaces are used to respond to an inference request based on a target inference device, and call the internal plug-in interface corresponding to the target inference device to convert the data to be inferred into service data; the internal plug-in interface includes an internal device interface, the external plug-in interface includes an external device interface, and the function plug-in includes a device plug-in for providing the internal device interface, wherein, the basic function module nestedly combines the internal device interfaces to generate an external device interface; when the inference request includes a device management request for the target inference device, the external device interface is used to call the internal device interface corresponding to the target inference device based on the device management request to manage the target inference device.

2. The inference framework system according to claim 1, characterized in that, the internal device interface includes a device binding sub-interface, a device status acquisition sub-interface, and a device memory operation sub-interface; wherein: the device binding sub-interface is used to obtain the identification information of the inference device; the device status acquisition sub-interface is used to obtain the device status of the inference device; the device memory operation sub-interface is used to operate the device memory.

3. The inference framework system according to claim 2, characterized in that, the internal plug-in interface includes an internal processing interface, the external plug-in interface includes an external processing interface, and the function plug-in includes a processing plug-in for providing the internal processing interface; wherein, the basic function module nestedly combines the internal processing interfaces of multiple inference devices to generate an external processing interface; when the inference request includes a pre-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the pre-processing request to convert the data to be inferred into the input data of the inference model; when the inference request includes a post-processing request, the external processing interface is used to call the internal processing interface corresponding to the target inference device based on the post-processing request to convert the output data of the inference model into service data.

4. The inference framework system according to claim 3, characterized in that, the internal processing interface includes a pre-processing sub-interface and a post-processing sub-interface; the pre-processing sub-interface is used to convert the data to be inferred into the input data of the inference model; the post-processing sub-interface is used to convert the output data of the inference model into service data.

5. The inference framework system according to claim 4, characterized in that, the internal plug-in interface includes an internal inference interface, the external plug-in interface includes an external inference interface, and the function plug-in includes an inference plug-in for providing the internal inference interface; wherein, The basic function module nests and combines the internal inference interfaces to generate an external inference interface. When the inference request includes a model inference request for the target inference device, the external inference interface is used to call the corresponding internal inference interface of the target inference device based on the model inference request, and convert the input data into output data based on the inference model.

6. The inference framework system according to claim 5, wherein, the internal inference interface includes internal inference interfaces corresponding to at least one inference model; the external inference interface is used to call the target internal inference interface corresponding to the target inference device based on the model inference request for the target inference device, and convert the input data into output data based on the target inference model; the inference request is used to call the target inference model in the at least one inference model, and the target internal inference interface is used to manage the inference process of the target inference model.

7. The inference framework system according to claim 5 or 6, wherein, the internal inference interface includes: a model parsing sub-interface, a model configuration sub-interface, a model inference sub-interface, and a result acquisition sub-interface; among them, the model parsing sub-interface is used to obtain a model file and perform parsing; the model configuration sub-interface is used to configure the parsed model to obtain a configured model; the model inference sub-interface is used to complete the model inference process based on the configured model; the result acquisition sub-interface is used to obtain the inference result of the model inference process.

8. The inference framework system according to claim 7, wherein, the basic function module is used to provide at least one business function; when the business function is called, at least one of the following is executed: when the business function is called by the internal inference interface, convert the input data into the output data; when the business function is called by the internal processing interface, convert the data to be inferred into the input data of the inference model; when the business function is called by the internal processing interface, convert the output data of the inference model into business data.

9. The inference framework system according to claim 3, wherein, the basic function module is used to provide an exchange data structure corresponding to each inference device; the inference request is used to convert the data to be inferred in the first data structure into business data in the second data structure; when the internal processing interface is called, it is further used to convert the data to be inferred in the first data structure into the data to be inferred in the third data structure; when the internal processing interface is called, it is further used to convert the business data in the fourth data structure into the business data in the second data structure; the third data structure and the fourth data structure are data structures supported by the target inference device.

10. The inference framework system according to claim 6, wherein, the basic function module is used to provide a model data structure corresponding to each inference model; When the internal inference interface is called, it is further used to convert the input data of the fifth data structure into the input data of the sixth data structure; the internal inference interface is used to convert the input data of the sixth data structure into the output data of the seventh data structure based on the target inference model; When the internal inference interface is called, it is further used to convert the output data of the seventh data structure into the service data of the eighth data structure.

11. The inference framework system according to claim 1, wherein, the basic function module is used to provide at least one of the following components: a cross-platform function component, a plugin management component, and a log management component, where: the cross-platform function component is used to provide a function library for each of the multiple system functions; the function library includes the system functions corresponding to at least one system platform; the plugin management component is used to manage the life cycle of the plugins in the plugin definition module; the log management component is used to monitor the life cycle change events of the plugins in the plugin definition module and generate corresponding logs based on the life cycle change events.

12. An inference method, wherein, applied to the inference framework system according to any one of claims 1 to 11, the inference framework system provides an external plugin interface; the method includes: receiving an inference request based on a target inference device; in response to the inference request, through the external plugin interface provided by the inference framework system, calling the internal plugin interface corresponding to the target inference device among the internal plugin interfaces of multiple inference devices provided by the inference framework system to convert the data to be inferred into service data.

13. The inference method according to claim 12, wherein, the external plugin interface includes an external device interface, an external processing interface, and an external inference interface; the step of, through the external plugin interface provided by the inference framework system, calling the internal plugin interface corresponding to the target inference device among the internal plugin interfaces of multiple inference devices provided by the inference framework system to convert the data to be inferred into service data includes: based on the external device interface provided by the inference framework system, calling the internal device interface corresponding to the target inference device in the inference framework system to load an inference model in the target inference device; based on the external processing interface provided by the inference framework system, calling the internal processing interface corresponding to the target inference device in the inference framework system to process the data to be inferred to obtain input data, and inputting the input data into the inference model; based on the external inference interface provided by the inference framework system, calling the internal inference interface corresponding to the target inference device in the inference framework system to convert the input data into output data by using the inference model; based on the external processing interface provided by the inference framework system, calling the internal processing interface corresponding to the target inference device in the inference framework system to convert the output data into the service data; Based on the external device interface provided by the inference framework system, call the internal device interface corresponding to the target inference device in the inference framework system, and release the inference model in the target inference device.

14. An inference device, characterized in that, applied to the inference framework system according to any one of claims 1 to 11, the device comprising: a receiving module, configured to receive an inference request based on a target inference device; an inference module, configured to, in response to the inference request, call the internal plug-in interface corresponding to the target inference device among the internal plug-in interfaces of multiple inference devices provided by the inference framework system through the external plug-in interface provided by the inference framework system to convert the data to be inferred into service data.

15. A computer device, comprising a memory and a processor, the memory storing a computer program that can be run on the processor, characterized in that, when the processor executes the program, the steps in the method according to any one of claims 12 to 13 are implemented.

16. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps in the method according to any one of claims 12 to 13 are implemented.