A processing method, device, computer device and storage medium for an inference service
By orchestrating the inference model in real-time in inference service processing, the problem that the inference model in the existing technology can only be applied to specific service scenarios is solved, the model reuse and efficient utilization of AI resources are achieved, cost is reduced and resource waste is avoided.
Patent Information
- Application Number
- CN202211019473.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-24
AI Technical Summary
In the prior art, inference models can only be applied to specific service scenarios, resulting in the inability to reuse the models, reducing the utilization rate of AI chips or AI graphics cards, increasing the cost of use, and causing waste of resources.
By obtaining the service type in the inference service request, determining the target service configuration in the preset initial service configuration, determining the target inference model in the pre-deployed initial inference model based on the model type configuration in the target service configuration, and orchestrating the target inference model to obtain the target service model to realize the processing of inference service.
The reuse of inference models is realized, the utilization rate of AI chips or AI graphics cards is improved, the cost of use is reduced, and the waste of resources is avoided.
Smart Images

Figure CN115358401B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, computer device and storage medium for processing inference services. Background Art
[0002] Model deployment refers to the process in which deployers deploy the required service models on the platform. This process mainly includes data compression of the models, improving concurrency, and improving compatibility with the platform. The service model consists of one or more specific AI (Artificial Intelligence) inference models; inference service refers to the process of using the target service model deployed on the platform to perform inference on the data to be inferred and obtain the inference result.
[0003] In the prior art, model deployment is carried out based on specific service scenarios. For example, when the service scenario is "character recognition in images" and this service scenario requires two inference models, during model deployment, these two inference models will be deployed on the platform according to the model parameters required by this service scenario to obtain the corresponding service model. That is, the service model that ultimately provides the inference service is directly obtained through deployment, and the two inference models corresponding to this service model can only be applied in this service scenario; this leads to the inability to reuse the deployed inference models. Since inference models generally need to rely on corresponding AI chips or AI graphics cards to improve inference performance, the utilization rate of AI chips or AI graphics cards is thus reduced. Since the cost of AI chips or AI graphics cards is relatively high, the usage cost is further increased; in addition, due to the inability to reuse, it is necessary to deploy the corresponding inference models again to obtain the service models required for other service scenarios, resulting in a waste of resources. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a method, device, computer device and storage medium for processing inference services.
[0005] In a first aspect, in one embodiment, the present invention provides a method for processing inference services, including:
[0006] Obtain the service type in the inference service request;
[0007] According to the service type, determine the target service configuration in the preset initial service configuration; according to the model type configuration in the target service configuration, determine the target inference model in the initially deployed initial inference models;
[0008] According to the model orchestration configuration in the target service configuration, orchestrate the target inference model to obtain the target service model;
[0009] Based on the target service model, perform inference on the data to be inferred corresponding to the inference service request to obtain an inference result.
[0010] In one embodiment, before the step of determining the target service configuration in the preset initial service configuration according to the service type, the processing method of the above inference service further includes:
[0011] Obtain the target configuration template in the preset initial configuration template and the target configuration parameters corresponding to the target configuration template;
[0012] According to the target configuration parameters, perform parameter setting on the target configuration template to obtain the initial service configuration.
[0013] In one embodiment, performing parameter setting on the target configuration template according to the target configuration parameters to obtain the initial service configuration includes:
[0014] According to the target configuration parameters, perform parameter setting on the target configuration template to obtain an intermediate service configuration;
[0015] According to the model type configuration in the intermediate service configuration, determine the intermediate inference model in the initially deployed initial inference model;
[0016] According to the model orchestration configuration in the intermediate service configuration, orchestrate the intermediate inference model to obtain an intermediate service model;
[0017] Obtain a training sample set, and train the intermediate service model according to the training sample set to obtain a trained intermediate service model;
[0018] According to the weight parameters of the trained intermediate service model, adjust the parameters of the model orchestration configuration in the intermediate service configuration to obtain the initial service configuration, and the model orchestration configuration includes model weight configuration.
[0019] In one embodiment, before the step of obtaining the target configuration template in the preset initial configuration template, the processing method of the above inference service further includes:
[0020] Obtain template construction data;
[0021] Render the initial configuration template according to the template construction data; the initial configuration template includes any one of a serial orchestration template, a branch orchestration template, a segmentation orchestration template, and an aggregation orchestration template.
[0022] In one embodiment, the model orchestration configuration includes model relationship configuration and model weight configuration; orchestrating the target inference model according to the model orchestration configuration in the target service configuration includes:
[0023] According to the model relationship configuration in the target service configuration, perform model combination on multiple target inference models;
[0024] Set the weights of multiple target inference models respectively according to the model weight configuration in the target service configuration.
[0025] In one embodiment, based on the target service model, perform inference on the data to be inferred corresponding to the inference service request to obtain an inference result, including:
[0026] Call the exposed interface of the input - end inference model in the target service model to receive the data to be inferred corresponding to the inference service request;
[0027] Based on the target service model, perform inference on the data to be inferred to obtain an inference result;
[0028] Call the exposed interface of the output - end inference model in the target service model to send the inference result.
[0029] In one embodiment, before the step of determining the target inference model in the initially deployed inference model according to the model type configuration in the target service configuration, the processing method of the above - mentioned inference service further includes:
[0030] Obtain model deployment data;
[0031] Deploy to obtain the initial inference model according to the model deployment data;
[0032] Set the exposed interface of the initial inference model.
[0033] In a second aspect, in one embodiment, the present invention provides a processing device for an inference service, including:
[0034] A type acquisition module, configured to acquire the service type in the inference service request;
[0035] A configuration determination module, configured to determine the target service configuration in the preset initial service configuration according to the service type; a model determination module, configured to determine the target inference model in the initially deployed inference model according to the model type configuration in the target service configuration;
[0036] A model orchestration module, configured to orchestrate the target inference model according to the model orchestration configuration in the target service configuration to obtain a target service model;
[0037] A data inference module, configured to perform inference on the data to be inferred corresponding to the inference service request based on the target service model to obtain an inference result.
[0038] In a third aspect, in one embodiment, the present invention provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the steps in the processing method of the inference service in any of the above embodiments.
[0039] In a fourth aspect, in one embodiment, the present invention provides a storage medium storing a computer program, and the computer program is loaded by a processor to execute the steps in the processing method of the inference service in any of the above embodiments.
[0040] Through the above processing method, device, computer device and storage medium of the inference service, the pre-deployed inference model is not restricted to any one service scenario, that is, the inference model is in an initialized state. When the target service configuration is determined according to the service type in the inference service request, the target inference model required can be arranged in real time according to the target service configuration, so as to obtain the corresponding target service model, and then the inference service is completed based on the target service model; since the target service model is arranged in real time according to the target service configuration, with the deployed target inference model as the basis and the target service configuration as the core variable for obtaining the target service model, the deployed inference model can be arranged by different service configurations, so as to obtain different service models, realizing the reuse of the inference model; improving the utilization rate of AI chips or AI graphics cards and reducing the usage cost; in addition, it also avoids the waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 It is a schematic diagram of the application scenario of the processing method of the inference service in one embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of the internal structure of the computer device in one embodiment of the present invention;
[0044] Figure 3 It is a schematic flowchart of the processing method of the inference service in one embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of the structure of the processing device of the inference service in one embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0047] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined. In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description". Any embodiment described as "exemplary" in this application is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in this application.
[0048] The processing method of the inference service in the embodiments of the present invention is applied to a processing device for the inference service, and the processing device for the inference service is provided in a computer device; the computer device can be a terminal, for example, a mobile phone or a tablet computer, and the computer device can also be a server, or a service cluster composed of multiple servers.
[0049] As Figure 1 shown, Figure 1 is a schematic diagram of the application scenario of the processing method of the inference service in the embodiments of the present invention. The application scenario of the processing method of the inference service in the embodiments of the present invention includes a computer device 100 (the processing device for the inference service is integrated in the computer device 100), and a computer-readable storage medium in the computer device 100 that runs the processing method of the inference service to execute the steps of the processing method of the inference service.
[0050] It can be understood that Figure 1The computer device in the application scenario of the processing method of the inference service shown, or the devices included in the computer device do not constitute a limitation to the embodiments of the present invention. That is, the number of devices, the types of devices included in the application scenario of the processing method of the inference service, or the number of devices and the types of devices included in each device do not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be regarded as equivalent replacements or derivatives of the technical solution claimed in the embodiments of the present invention.
[0051] In the embodiments of the present invention, the computer device 100 can be an independent device, or a network or cluster of devices composed of devices. For example, the computer device 100 described in the embodiments of the present invention includes, but is not limited to, a computer, a network host, a single network device, a set of multiple network devices, or a cloud device composed of multiple devices. Among them, the cloud device is composed of a large number of computers or network devices based on cloud computing.
[0052] Those skilled in the art can understand that Figure 1 the application scenario shown is only one application scenario corresponding to the technical solution of the present invention, and does not constitute a limitation to the application scenario of the technical solution of the present invention. Other application scenarios can also include more or fewer computer devices than Figure 1 shown, or the network connection relationship of computer devices. For example Figure 1 only 1 computer device is shown in . It can be understood that the scenario of the processing method of the inference service can also include one or more other computer devices, which are not specifically limited here; the computer device 100 can also include a memory for storing information related to the processing method of the inference service.
[0053] In addition, in the application scenario of the processing method of the inference service in the embodiments of the present invention, the computer device 100 can be provided with a display device, or the computer device 100 is not provided with a display device and is communicatively connected to an external display device 200. The display device 200 is used to output the result of the execution of the processing method of the inference service in the computer device. The computer device 100 can access the background database 300 (the background database 300 can be the local memory of the computer device 100, and the background database 300 can also be set in the cloud), and the background database 300 stores information related to the processing method of the inference service.
[0054] It should be noted that Figure 1 the application scenario of the processing method of the inference service shown is only an example. The application scenario of the processing method of the inference service described in the embodiments of the present invention is to more clearly illustrate the technical solution of the embodiments of the present invention, and does not constitute a limitation to the technical solution provided by the embodiments of the present invention.
[0055] Such as Figure 2As shown, it shows the structure of the computer device involved in the present invention. Specifically:
[0056] The computer device may include components such as a processor 201 with one or more processing cores, a memory 202 with one or more computer-readable storage media, a power supply 203, and an input unit 204. Those skilled in the art can understand that Figure 2 the structure of the computer device shown in does not limit the computer device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0057] The processor 201 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 202, and by calling the data stored in the memory 202, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 201 may include one or more processing cores; preferably, the processor 201 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and computer programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 201.
[0058] The memory 202 can be used to store software programs and modules. The processor 201 executes various functional applications and data processing by running the software programs and modules stored in the memory 202. The memory 202 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, computer programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the server. In addition, the memory 202 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 202 may also include a memory controller to provide the processor 201 with access to the memory 202.
[0059] The computer device also includes a power supply 203 that powers each component. Preferably, the power supply 203 can be logically connected to the processor 201 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 203 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0060] The computer device may further include an input unit 204, which may be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0061] Based on the application scenario of the processing method of the above-mentioned inference service, embodiments of the processing method of the inference service are proposed.
[0062] In a first aspect, as Figure 3 shown, in an embodiment, the present invention provides a processing method for an inference service, including:
[0063] Step 301, obtain the service type in the inference service request;
[0064] Among them, this embodiment is mainly used to provide an online inference service, that is, the execution subject of the processing method of the inference service in this embodiment is the service provider, and the service provider receives the inference service request sent by the requestor;
[0065] Among them, the service type is mainly used to determine the service scenario, so as to determine the service configuration corresponding to the service scenario; the requestor and the service provider can pre-agree on the representation form of the service type, so that when the service provider obtains the service type sent by the requestor, it can recognize its meaning; specifically, the service provider can pre-provide the service type corresponding to each initial service configuration, and the requestor can send selection information for a certain service type, so that the service provider determines the corresponding service type as the service type specified in the service processing request according to the selection information;
[0066] Among them, the service provider and the requestor can send and receive inference service requests through the HTTP protocol (HyperText Transfer Protocol, which is an application layer transport protocol based on the TCP protocol. Simply put, it is a rule for data transmission between the client and the server), that is, the inference service request is an http request (the http request includes an http request line, an http request header, and an http request body). In this http request, the data corresponding to the service type is stored in the http request body;
[0067] Step 302, determine the target service configuration in the preset initial service configuration according to the service type;
[0068] Among them, the service type and the service configuration are corresponding, and the service type can be used as the code identifier of the service configuration. Specifically, a service configuration library can be pre-constructed, which contains multiple initial service configurations and their corresponding initial service types. Therefore, the initial service configuration corresponding to the service type in the inference service request can be matched in the service configuration library, and then the initial service configuration corresponding to the initial service type is set as the target service configuration required for this inference service. For example, if the service scenario represented by the obtained service type is "contour detection of an image", the target service configuration corresponding to "contour detection of an image" can be determined.
[0069] Among them, the service configuration includes a model type configuration and a model orchestration configuration. The model type configuration mainly specifies the model type of the inference model to be used, while the model orchestration configuration mainly specifies the various parameters required in the process of implementing the inference model to the service model.
[0070] Step 303: Determine the target inference model in the pre-deployed initial inference models according to the model type configuration in the target service configuration.
[0071] Among them, the model type in the model type configuration and the inference model are corresponding, and the model type can be used as the code identifier of the inference model. Specifically, an inference model library can be pre-constructed, which contains multiple initial inference models and their corresponding initial model types (it should be noted that the inference model library is only used to explain an implementation method of this embodiment, rather than storing each inference model in the inference model library. In fact, each inference model has been deployed on the platform of the service provider). Therefore, the initial inference model corresponding to each model type in the model type configuration can be matched in the inference model library, and then the initial inference model corresponding to the initial model type is set as the corresponding target inference type. For example, if the target service configuration of "contour detection of an image" is obtained, the corresponding model type configuration is the target inference model required to implement "contour detection of an image", such as inference model A.
[0072] Step 304: Orchestrate the target inference model according to the model orchestration configuration in the target service configuration to obtain the target service model.
[0073] Among them, the determined target inference model is only in an initialized state, and its weight parameters are in an "editable" state. When not edited, the target inference model cannot execute properly or the execution effect is far from the expectation (this situation mainly refers to the existence of an initialized weight parameter). Therefore, when the model orchestration configuration is determined, the target inference model can be parameterized according to the weight parameters therein, so as to obtain the corresponding target service model. Specifically, for example, if the target inference model only includes inference model A, then the model orchestration configuration only contains the weight parameters of inference model A, and then inference model A is parameterized according to the weight parameters to obtain inference model A with weight parameters, that is, the target service model. In this example, the target service model consists of only one inference model. Of course, in other embodiments, the target inference model can also be composed of two or more inference models, specifically depending on the corresponding service scenario;
[0074] Step 305: Based on the target service model, infer the data to be inferred corresponding to the inference service request to obtain an inference result;
[0075] Among them, after obtaining the target service model, the inference service can be performed. For example, taking the service scenario of "contour detection of an image" mentioned above as an example, the data to be inferred is the image to be detected, and the inference result is the contour image corresponding to the image to be detected;
[0076] Among them, the data to be inferred can be directly included in the above-mentioned inference service request (in this method, it is necessary to parse the inference service request to obtain the data to be inferred), or stored at an address, and this address is included in the inference service request sent by the requester, so as to reduce the data volume of the inference service request, and when the communication link between this address and the service provider is relatively close, transmission resources can be saved.
[0077] Through the above processing method of the inference service, the pre-deployed inference model is not restricted to any service scenario, that is, the inference model is in an initialized state. When the target service configuration is determined according to the service type in the inference service request, the required target inference model can be orchestrated in real time according to the target service configuration, so as to obtain the corresponding target service model, and then the inference service is completed based on the target service model. Since the target service model is obtained by real-time orchestration according to the target service configuration, with the deployed target inference model as the basis and the target service configuration as the core variable for obtaining the target service model, the deployed inference model can be orchestrated by different service configurations, so as to obtain different service models, realizing the reuse of the inference model; improving the utilization rate of AI chips or AI graphics cards and reducing the usage cost; in addition, it also avoids waste of resources.
[0078] In one embodiment, before the step of determining the target service configuration in the preset initial service configuration according to the service type, the processing method of the inference service further includes:
[0079] Obtain the target configuration template in the preset initial configuration template and the target configuration parameters corresponding to the target configuration template;
[0080] Among them, service providers can build the initial service configuration required for subsequent inference services through a visual management interface. Specifically, there are multiple initial configuration templates preset in the visual management interface, and each initial configuration template adopts the logical method of DAG (Directed Acyclic Graph). Service providers can select the corresponding initial configuration template according to the required service scenario. For example, by inputting a click instruction for a certain initial configuration template in the visual management interface through the mouse, the initial configuration template can be set as the target configuration template. Similarly, service providers can input corresponding target configuration parameters according to the required service scenario. Specifically, they can directly input the complete target configuration parameters through the keyboard and / or mouse, or select one of the multiple typical configuration parameters provided by the interface as the target configuration parameter, which is not limited here;
[0081] Set parameters for the target configuration template according to the target configuration parameters to obtain the initial service configuration;
[0082] Among them, service providers usually input corresponding target configuration parameters based on the selected target configuration template. A DAG is a graph structure composed of vertices and directed edges. In this graph, starting from a selected vertex v and searching along the ordered edges, it is ultimately impossible to return to vertex v, that is, no cycle will be formed; the target configuration parameters can specifically characterize the meaning of each node including vertices in the DAG, as well as the directed edge relationships between the nodes; for example, the ultimately expected target service model consists of inference model A, inference model B, inference model C, and inference model D, and the implementation process includes: inputting the data to be inferred into inference model A, and then after obtaining the first result output by inference model A, inputting the first result into inference model B and inference model C respectively, and then obtaining the second result output by inference model B and the third result output by inference model C respectively. Finally, inputting the second result and the third result into inference model D, and then obtaining the inference result output by inference model D; then in the DAG, there are four corresponding nodes, namely node A, node B, node C, and node D. These four nodes respectively represent inference model A (including the weight parameters of inference model A), inference model B (including the weight parameters of inference model B), inference model C (including the weight parameters of inference model C), and inference model D (including the weight parameters of inference model D). Then node A, as the vertex, is connected to node B and node C respectively, and node B and node C are respectively connected to node D. And the DAG containing the above information is used as the required initial service configuration. After subsequently determining this initial service configuration as the target service configuration, the corresponding target service model can be obtained according to the target service configuration; the essence of the service provider is to construct a DAG with the above properties. Specifically, the target configuration template is a blank DAG. The blank DAG has a certain number of nodes and certain directed edge relationships between the nodes. The service provider adjusts the number of nodes of the blank DAG, as well as the directed edge relationships between the nodes, and sets the meaning represented by each node, and finally constructs an initial service configuration.
[0083] Among them, when service providers perform parameter settings, they also need to add a custom business logic intermediate price in the responsibility chain mode (definition: to avoid the coupling of the request sender with multiple request handlers, all request handlers are connected into a chain by the previous object remembering the reference of its next object; when a request occurs, the request can be passed along this chain until an object processes it), so as to ensure that the service model can be reliably obtained according to the service configuration subsequently.
[0084] Among them, when the service provider's personnel set the weight parameters of each node, the input weight parameters can be pre-trained. That is, the service provider's personnel can deploy the target service model on other platforms or systems, then train it, and finally obtain the trained target service model and the weight parameters corresponding to the trained target service model. Then, the weight parameters of each node can be directly set according to the weight parameters.
[0085] By means of a visual management interface and developing a corresponding initial configuration template, the creation of the initial service configuration can be made more efficient in the subsequent process, saving labor costs.
[0086] In one embodiment, according to the target configuration parameters, parameter setting is performed on the target configuration template to obtain an initial service configuration, including:
[0087] According to the target configuration parameters, parameter setting is performed on the target configuration template to obtain an intermediate service configuration;
[0088] As mentioned in the above embodiment, the weight parameters in the target configuration parameters can be pre-trained; while in this embodiment, the weight parameters in the target configuration parameters can be untrained. Therefore, it is necessary to call the corresponding inference model on this platform for training. Specifically, the service provider's personnel can input an initial weight parameter for parameter setting to obtain an intermediate service configuration (in this embodiment, the model type and the directed edge relationship are not emphasized, and the specific details can refer to the above embodiment);
[0089] According to the model type configuration in the intermediate service configuration, determine the intermediate inference model in the initially deployed initial inference model;
[0090] According to the model orchestration configuration in the intermediate service configuration, orchestrate the intermediate inference model to obtain an intermediate service model;
[0091] Among them, the specific process of obtaining the corresponding intermediate service model according to the intermediate service configuration can refer to the specific process of "obtaining the corresponding target service model according to the target service configuration" in the above embodiment. The essential steps of the two are basically the same and will not be elaborated here;
[0092] Obtain a training sample set, and train the intermediate service model according to the training sample set to obtain a trained intermediate service model;
[0093] Among them, the training sample set can adopt publicly available samples, and the purpose of training is to make the output of the model meet the expected requirements;
[0094] According to the weight parameters of the trained intermediate service model, adjust the parameters of the intermediate service configuration to obtain an initial service configuration;
[0095] Among them, the parameter adjustment of the intermediate service configuration is mainly for the weight parameter part in the intermediate service configuration. After adjusting the weight parameters, the weight parameters in the intermediate service configuration are the trained weight parameters. Therefore, the intermediate service configuration can be directly determined as the corresponding initial service configuration.
[0096] In one embodiment, obtaining a training sample set, training an intermediate service model according to the training sample set to obtain a trained intermediate service model, includes:
[0097] Obtaining a training sample set, where the training sample set includes multiple training samples, and each training sample includes training data to be inferred and training inference results;
[0098] Among them, the training inference results are obtained by manually processing the training data to be inferred;
[0099] Obtaining a training sample, using the training data to be inferred as the input of the intermediate service model, and using the training inference result as the expected output of the intermediate service model, and training the intermediate service model, that is, completing one training;
[0100] Determining the comparison result between the actual output and the expected output of the intermediate service model. If the comparison result does not meet the requirements, updating the model parameters of the intermediate service model according to the comparison result;
[0101] Obtaining the next training sample, and then re-entering the step of using the training data to be inferred as the input of the intermediate service model and using the training inference result as the expected output of the intermediate service model to train the intermediate service model, until the obtained comparison result meets the requirements, and stopping the training to obtain a trained intermediate service model.
[0102] In one embodiment, determining the comparison result between the actual output and the expected output of the intermediate service model. If the comparison result does not meet the requirements, updating the model parameters of the intermediate service model according to the comparison result, includes:
[0103] Determining the comparison difference between the actual output and the expected output of the intermediate service model, and calculating a loss value according to the comparison difference;
[0104] If the loss value does not meet the preset convergence condition, updating the model parameters of the intermediate service model according to the loss value.
[0105] In one embodiment, before the step of obtaining the target configuration template in the preset initial configuration template, the above-mentioned processing method of the inference service further includes:
[0106] Obtaining template construction data;
[0107] Among them, as mentioned above, service providers can obtain the corresponding initial service configuration by selecting a configuration template in the visual management interface. Therefore, it is necessary to construct the configuration template in the visual management interface before that. Specifically, since the configuration template is essentially a blank DAG, service providers can construct the configuration template by means of primitive editing. The template construction data refers to various instructions or parameters input by service providers during the construction process.
[0108] Render the initial configuration template according to the template construction data.
[0109] Among them, the visual management interface renders in real time according to various instructions or parameters input by service providers in the corresponding editing interface. When service providers complete the input of all instructions or parameters, the corresponding initial configuration template can be rendered.
[0110] Among them, the initial configuration template includes any one of a serial orchestration template, a branch orchestration template, a splitting orchestration template, and an aggregation orchestration template.
[0111] In one embodiment, the model orchestration configuration includes model relationship configuration and model weight configuration. Orchestrating the target inference model according to the model orchestration configuration in the target service configuration includes:
[0112] Combining multiple target inference models according to the model relationship configuration in the target service configuration.
[0113] Among them, as mentioned above, the target service configuration exists in the form of a DAG. Therefore, when the corresponding target service model consists of multiple inference models, there will be multiple nodes in the DAG. In addition to setting the meaning of each node (including the model type (i.e., model type configuration) and weight parameters (i.e., model weight configuration in this embodiment)), it is also necessary to set the directed edge relationship between multiple nodes for multiple nodes, which is the model relationship configuration in this embodiment. If the target service model consists of only one inference model, there will be only one node in the DAG, so there is no need to consider the directed edge relationship between nodes.
[0114] Set the weights of multiple target inference models respectively according to the model weight configuration in the target service configuration.
[0115] Among them, the model weight configuration is mainly used to set the weight parameters of the inference model. For specific details, reference can be made to the above embodiments and will not be elaborated here.
[0116] In one embodiment, based on the target service model, infer the data to be inferred corresponding to the inference service request to obtain an inference result, including:
[0117] Invoke the exposed interface of the inference model at the input end in the target service model to receive the data to be inferred corresponding to the inference service request;
[0118] Among them, the exposed interface is an HTTP standard inference interface. There is a corresponding exposed interface for each initial inference model. When the target service model consists of multiple inference models, the inference model used to process the data to be inferred first is the inference model at the input end. For example, the target service model finally expected to be obtained consists of inference models A, B, C, and D, and the implementation process includes: inputting the data to be inferred into inference model A, and then after obtaining the first result output by inference model A, inputting the first result into inference models B and C respectively, and then obtaining the second result output by inference model B and the third result output by inference model C respectively. Finally, inputting the second result and the third result into inference model D, and then obtaining the inference result output by inference model D. Then inference model A is the inference model at the input end. After obtaining this target service model, the exposed interface of inference model A can be invoked to directly receive the data to be inferred sent by the requester;
[0119] Based on the target service model, infer the data to be inferred to obtain an inference result;
[0120] Invoke the exposed interface of the inference model at the output end in the target service model to send the inference result;
[0121] Among them, similarly, when the target service model consists of multiple inference models, the inference model used to process the data to be inferred finally and obtain the inference result is the inference model at the output end. Taking the example in the previous step as an example, inference model D is the inference model at the output end, which is used to send the inference result to the requester;
[0122] Among them, it should be noted that when the target service model consists of one inference model, then this inference model serves as both the inference model at the input end and the inference model at the output end. The reception of the data to be inferred and the sending of the inference result are both through the exposed interface of this inference model;
[0123] Among them, when there is an exposed interface, the data to be processed does not need to be included in the inference service request, nor does it need to be stored at a certain address.
[0124] In one embodiment, before the step of determining the target inference model in the initially deployed initial inference models according to the model type configuration in the target service configuration, the above-mentioned processing method of the inference service further includes:
[0125] Obtain model deployment data;
[0126] Deploy the initial inference model according to the model deployment data;
[0127] Among them, model deployment mainly follows requirements such as data compression, improving concurrency, and enhancing compatibility with the platform, and assembles the corresponding model on the platform. The specific process of model deployment can refer to the existing technology and will not be elaborated here.
[0128] Set the exposed interface of the initial inference model.
[0129] Among them, setting the exposed interface is used to implement the reception of subsequent data to be inferred and the sending of inference results. For specific details, reference can be made to the above embodiments and will not be elaborated here.
[0130] In a second aspect, as Figure 4 shown, in one embodiment, the present invention provides a processing device for inference services, including:
[0131] A type acquisition module 401, configured to acquire the service type in the inference service request.
[0132] A configuration determination module 402, configured to determine the target service configuration in the preset initial service configuration according to the service type.
[0133] A model determination module 403, configured to determine the target inference model in the pre-deployed initial inference model according to the model type configuration in the target service configuration.
[0134] A model orchestration module 404, configured to orchestrate the target inference model according to the model orchestration configuration in the target service configuration to obtain a target service model.
[0135] A data inference module 405, configured to perform inference on the data to be inferred corresponding to the inference service request based on the target service model to obtain an inference result.
[0136] Through the above-mentioned processing device for inference services, the pre-deployed inference model is not restricted to any one service scenario, that is, the inference model is in an initialized state. When the target service configuration is determined according to the service type in the inference service request, the target inference model required can be orchestrated in real time according to the target service configuration, so as to obtain the corresponding target service model, and then complete the inference service based on the target service model. Since the target service model is obtained by real-time orchestration according to the target service configuration, with the deployed target inference model as the basis and the target service configuration as the core variable for obtaining the target service model, the deployed inference model can be orchestrated by different service configurations to obtain different service models, realizing the reuse of the inference model, improving the utilization rate of AI chips or AI graphics cards, reducing the usage cost, and also avoiding waste of resources.
[0137] In one embodiment, the above-mentioned processing device for inference services further includes:
[0138] A configuration setting module, before the step of determining a target service configuration in a preset initial service configuration according to the service type, obtains a target configuration template in a preset initial configuration template and target configuration parameters corresponding to the target configuration template; and performs parameter setting on the target configuration template according to the target configuration parameters to obtain an initial service configuration.
[0139] In one embodiment, the configuration setting module is specifically configured to perform parameter setting on the target configuration template according to the target configuration parameters to obtain an intermediate service configuration; determine an intermediate inference model in a pre-deployed initial inference model according to the model type configuration in the intermediate service configuration; perform orchestration on the intermediate inference model according to the model orchestration configuration in the intermediate service configuration to obtain an intermediate service model; obtain a training sample set, and train the intermediate service model according to the training sample set to obtain a trained intermediate service model; and perform parameter adjustment on the model orchestration configuration in the intermediate service configuration according to the weight parameters of the trained intermediate service model to obtain an initial service configuration, where the model orchestration configuration includes a model weight configuration.
[0140] In one embodiment, the processing device for the inference service further includes:
[0141] A template setting module, before the step of obtaining a target configuration template in a preset initial configuration template, obtains template construction data; and renders an initial configuration template according to the template construction data; the initial configuration template includes any one of a serial orchestration template, a branch orchestration template, a segmentation orchestration template, and an aggregation orchestration template.
[0142] In one embodiment, the model orchestration configuration includes a model relationship configuration and a model weight configuration; the model orchestration module is specifically configured to perform model combination on a plurality of target inference models according to the model relationship configuration in the target service configuration; and perform weight setting on the plurality of target inference models respectively according to the model weight configuration in the target service configuration.
[0143] In one embodiment, the data inference module is specifically configured to call an exposed interface of an input-end inference model in the target service model, receive data to be inferred corresponding to an inference service request; perform inference on the data to be inferred based on the target service model to obtain an inference result; and call an exposed interface of an output-end inference model in the target service model to send the inference result.
[0144] In one embodiment, the processing device for the inference service further includes:
[0145] The model deployment module is used to obtain model deployment data before determining the target inference model in the pre-deployed initial inference model according to the model type configuration in the target service configuration; deploy the initial inference model according to the model deployment data; and set the exposed interface of the initial inference model.
[0146] In a third aspect, in one embodiment, the present invention provides a computer device. Specifically, in this embodiment, the processor 201 in the computer device will load the executable files corresponding to the processes of one or more computer programs into the memory 202 according to the following instructions, and the processor 201 will run the computer programs stored in the memory 202 to perform the following steps:
[0147] Obtain the service type in the inference service request;
[0148] According to the service type, determine the target service configuration in the preset initial service configuration; according to the model type configuration in the target service configuration, determine the target inference model in the pre-deployed initial inference model;
[0149] Orchestrate the target inference model according to the model orchestration configuration in the target service configuration to obtain a target service model;
[0150] Based on the target service model, perform inference on the data to be inferred corresponding to the inference service request to obtain an inference result.
[0151] Through the above computer device, the pre-deployed inference model is not restricted to any one service scenario, that is, the inference model is in an initialized state. When the target service configuration is determined according to the service type in the inference service request, the required target inference model can be orchestrated in real time according to the target service configuration, so as to obtain the corresponding target service model, and then the inference service can be completed based on the target service model; since the target service model is orchestrated in real time according to the target service configuration, with the deployed target inference model as the basis and the target service configuration as the core variable for obtaining the target service model, the deployed inference model can be orchestrated by different service configurations to obtain different service models, realizing the reuse of the inference model; improving the utilization rate of AI chips or AI graphics cards and reducing the usage cost; in addition, it also avoids waste of resources.
[0152] Those of ordinary skill in the art can understand that all or part of the steps in any of the above methods can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0153] Fourthly, in one embodiment, the present invention provides a storage medium storing multiple computer programs that can be loaded by a processor to execute the following steps:
[0154] Obtain the service type in the inference service request;
[0155] According to the service type, determine the target service configuration in the preset initial service configuration; according to the model type configuration in the target service configuration, determine the target inference model in the pre-deployed initial inference models;
[0156] According to the model orchestration configuration in the target service configuration, orchestrate the target inference model to obtain a target service model;
[0157] Based on the target service model, perform inference on the data to be inferred corresponding to the inference service request to obtain an inference result.
[0158] Through the above storage medium, the pre-deployed inference model is not restricted to any one service scenario, that is, the inference model is in an initialized state. When the target service configuration is determined according to the service type in the inference service request, the required target inference model can be orchestrated in real time according to the target service configuration, so as to obtain the corresponding target service model, and then complete the inference service based on the target service model; since the target service model is obtained by real-time orchestration according to the target service configuration, with the deployed target inference model as the basis and the target service configuration as the core variable for obtaining the target service model, the deployed inference model can be orchestrated by different service configurations to obtain different service models, realizing the reuse of the inference model; improving the utilization rate of AI chips or AI graphics cards and reducing the usage cost; in addition, it also avoids waste of resources.
[0159] Those of ordinary skill in the art can understand that any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0160] Since the computer program stored in the storage medium can execute the steps in the processing method of the inference service in any one of the embodiments provided by the present invention, the beneficial effects achievable by the processing method of the inference service in any one of the embodiments provided by the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0161] For the specific implementation of each of the above operations, refer to the previous embodiments and will not be elaborated here.
[0162] In the above embodiments, the descriptions of the various embodiments have their respective emphases. For parts not elaborated in a certain embodiment, refer to the detailed descriptions of other embodiments above. This will not be elaborated here.
[0163] The above has introduced in detail a processing method, device, computer device, and storage medium for an inference service provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0164] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
Claims
1. A processing method for an inference service, characterized in that, Including: Obtain the service type in the inference service request; Determine the target service configuration in the preset initial service configuration according to the service type; Determine the target inference model in the pre-deployed initial inference model according to the model type configuration in the target service configuration; Orchestrate the target inference model according to the model orchestration configuration in the target service configuration to obtain a target service model; Based on the target service model, perform inference on the data to be inferred corresponding to the inference service request to obtain an inference result; Before the step of determining the target service configuration in the preset initial service configuration according to the service type, it further includes: Obtain the target configuration template in the preset initial configuration template and the target configuration parameters corresponding to the target configuration template; perform parameter setting on the target configuration template according to the target configuration parameters to obtain the initial service configuration; The step of performing parameter setting on the target configuration template according to the target configuration parameters to obtain the initial service configuration includes: Perform parameter setting on the target configuration template according to the target configuration parameters to obtain an intermediate service configuration; determine the intermediate inference model in the pre-deployed initial inference model according to the model type configuration in the intermediate service configuration; orchestrate the intermediate inference model according to the model orchestration configuration in the intermediate service configuration to obtain an intermediate service model; obtain a training sample set, train the intermediate service model according to the training sample set to obtain a trained intermediate service model; perform parameter adjustment on the model orchestration configuration in the intermediate service configuration according to the weight parameters of the trained intermediate service model to obtain the initial service configuration, and the model orchestration configuration includes model weight configuration; The model orchestration configuration includes model relationship configuration and model weight configuration; the step of orchestrating the target inference model according to the model orchestration configuration in the target service configuration includes: Perform model combination on multiple target inference models according to the model relationship configuration in the target service configuration; perform weight setting on multiple target inference models respectively according to the model weight configuration in the target service configuration.
2. The processing method of the inference service according to claim 1, wherein Before the step of obtaining the target configuration template in the preset initial configuration template, it further includes: Obtain template construction data; Render the initial configuration template according to the template construction data; the initial configuration template includes any one of a serial orchestration template, a branch orchestration template, a segmentation orchestration template, and an aggregation orchestration template.
3. The processing method of the inference service according to claim 1, wherein The step of performing inference on the data to be inferred corresponding to the inference service request based on the target service model to obtain an inference result includes: Invoke the exposed interface of the input-end inference model in the target service model to receive the data to be inferred corresponding to the inference service request; Perform inference on the data to be inferred based on the target service model to obtain the inference result; Invoke the exposed interface of the output-end inference model in the target service model to send the inference result.
4. The processing method of the inference service according to claim 3, characterized in that, Before the step of determining a target inference model in a pre-deployed initial inference model according to the model type configuration in the target service configuration, the following steps are further included: Obtain model deployment data; Deploy the initial inference model according to the model deployment data; Set the exposure interface of the initial inference model.
5. A processing device for an inference service, characterized in that, It includes: A type acquisition module for acquiring the service type in an inference service request; A configuration determination module for determining a target service configuration in a preset initial service configuration according to the service type; A model determination module for determining a target inference model in a pre-deployed initial inference model according to the model type configuration in the target service configuration; A model orchestration module for orchestrating the target inference model according to the model orchestration configuration in the target service configuration to obtain a target service model; A data inference module for inferring the data to be inferred corresponding to the inference service request based on the target service model to obtain an inference result; A configuration setting module for, before the step of determining a target service configuration in a preset initial service configuration according to the service type, acquiring a target configuration template in a preset initial configuration template and target configuration parameters corresponding to the target configuration template, and performing parameter setting on the target configuration template according to the target configuration parameters to obtain the initial service configuration; The configuration setting module is specifically configured to perform parameter setting on the target configuration template according to the target configuration parameters to obtain an intermediate service configuration, determine an intermediate inference model in a pre-deployed initial inference model according to the model type configuration in the intermediate service configuration, orchestrate the intermediate inference model according to the model orchestration configuration in the intermediate service configuration to obtain an intermediate service model, acquire a training sample set, train the intermediate service model according to the training sample set to obtain a trained intermediate service model, and perform parameter adjustment on the model orchestration configuration in the intermediate service configuration according to the weight parameters of the trained intermediate service model to obtain the initial service configuration, and the model orchestration configuration includes a model weight configuration; The model orchestration configuration includes a model relationship configuration and a model weight configuration; the model orchestration module is specifically configured to perform model combination on multiple target inference models according to the model relationship configuration in the target service configuration, and perform weight setting on multiple target inference models respectively according to the model weight configuration in the target service configuration.
6. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the steps in the processing method of the inference service according to any one of claims 1 to 5.
7. A storage medium, characterized in that, The storage medium stores a computer program, and the computer program is loaded by a processor to execute the steps in the processing method of the inference service according to any one of claims 1 to 5.
Citation Information
Patent Citations
Model deployment method and device, target monitoring method and device, equipment and system
CN110808881A
Inference service deployment method and device, computer equipment and storage medium
CN114819160A