Model deployment method, identification method, device and electronic equipment

By building service images and generating interfaces through the model deployment platform, the complexity of the neural network model deployment process is solved, enabling rapid deployment and verification, reducing user technical requirements, and improving efficiency.

CN114253556BActive Publication Date: 2026-03-20QINGDAO HAIER TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, the deployment process of neural network models is cumbersome and costly, which is not user-friendly for ordinary users. In particular, building services on one's own is time-consuming, and models for different platforms require specific training, which increases the difficulty and threshold of deployment.

Method used

This paper provides a model deployment method that builds a service image through a model deployment platform, obtains inference instances using deployment tools, and generates web page links, input interfaces, and output interfaces, simplifying the operation process and reducing the technical requirements for users.

Benefits of technology

It enables rapid deployment and validation of neural network models, lowers the deployment threshold, improves efficiency, and allows ordinary users to easily deploy and validate models using the model deployment platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253556B_ABST
    Figure CN114253556B_ABST
Patent Text Reader

Abstract

The application provides a model deployment method, a recognition method, a device and electronic equipment, comprising: constructing a service image of a neural network model according to a deployment file of the neural network model; obtaining an inference instance by using a deployment tool of a model deployment platform according to the service image; and generating a web link, an input interface and an output interface of the neural network model according to the inference instance. The model deployment method, the recognition method, the device and the electronic equipment provided by the application can deploy the neural network model according to a deployment file provided by a user by using the model deployment platform, and generate corresponding web links, input interfaces and output interfaces. The user only needs to understand the capability of the model deployment platform and learn to write a configuration file, and can quickly deploy the model by using the model deployment platform, so that the operation is simplified, the model deployment efficiency is improved, the threshold of model deployment is reduced, and the user is friendly.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a model deployment method, a recognition method, a device and an electronic device. BACKGROUND

[0002] The development of deep learning technology is constantly impacting more and more fields, and personnel in various fields are committed to optimizing their own solutions and user experience using deep learning technology. A large number of technical documents and open source materials can help developers obtain a satisfactory deep learning model. As deep learning models are increasingly easy to obtain, after training a neural network model, it is necessary to deploy the neural network model to obtain actual on-site use feedback in order to test the effect on hardware. The technology required for model deployment is different from the training of the model, which will hinder the rapid verification of the model effect to some extent.

[0003] The deployment and testing of a neural network model is a very tedious task. Current model deployment technologies mainly include self-built services or the use of service platform systems. Among them, self-built services are time-consuming and costly, and each service platform system requires the model to meet the specifications of different platforms, and even the model must be obtained through a specific platform.

[0004] The above solutions are targeted at computer program developers, and users need to master relevant technologies and understand the deployment environment while handling exceptions that occur during the deployment process. This is not user-friendly for ordinary users without development experience. SUMMARY

[0005] To solve the problems in the prior art, the embodiments of the present application provide a model deployment method, a recognition method, a device and an electronic device.

[0006] The present application provides a model deployment method, comprising: constructing a service image of a neural network model according to a deployment file of the neural network model;

[0007] According to the service image, an inference instance is obtained using a deployment tool of the model deployment platform.

[0008] According to the inference instance, a web link, an input interface and an output interface of the neural network model are generated.

[0009] The present application provides a recognition method, comprising:

[0010] Receiving a to-be-processed object input from an input interface, the to-be-processed object being image data or voice data;

[0011] The neural network model is used for identifying the to-be-processed object.

[0012] The output interface is used for outputting the identification result.

[0013] The inference instance, the webpage link, the input interface and the output interface are determined based on the model deployment method.

[0014] The application further provides a model deployment device, comprising: a construction module, configured to construct a service image of a neural network model according to a deployment file of the neural network model; the neural network model is an image recognition model or a voice recognition model;

[0015] An acquisition module is configured to acquire an inference instance by using a deployment tool of the model deployment platform according to the service image.

[0016] A first generation module is configured to generate a webpage link, an input interface and an output interface of the neural network model according to the inference instance.

[0017] The application further provides a recognition device, comprising:

[0018] A receiving module is configured to receive a to-be-processed object input from an input interface; the to-be-processed object is image data or voice data.

[0019] A second generation module is configured to identify the to-be-processed object by using an inference instance corresponding to a neural network model based on a webpage link, and generate an identification result; the neural network model is used for identifying the to-be-processed object.

[0020] An output module is configured to output the identification result by using an output interface.

[0021] The inference instance, the webpage link, the input interface and the output interface are determined based on the model deployment method.

[0022] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor; when the processor executes the program, the steps of the model deployment method or the recognition method are realized.

[0023] The application further provides a non-transitory computer readable storage medium, which stores a computer program; when the computer program is executed by a processor, the steps of the model deployment method or the recognition method are realized.

[0024] The model deployment method, the identification method, the device and the electronic equipment provided by the application, the model deployment platform deploys the neural network model according to the deployment file provided by the user, generates corresponding webpage links, input interfaces and output interfaces, the user only needs to understand the ability of the model deployment platform and learn to write the configuration file, and then the model can be quickly deployed by using the model deployment platform, the operation is simplified, the model deployment efficiency is improved, the threshold of model deployment is reduced, and the user is friendly. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0026] Figure 1 It is a flowchart of the model deployment method provided by the application;

[0027] Figure 2 It is a structural diagram of the model deployment platform provided by the application;

[0028] Figure 3 It is a flowchart of the identification method provided by the application;

[0029] Figure 4 It is a structural diagram of the model deployment device provided by the application;

[0030] Figure 5 It is a structural diagram of the identification device provided by the application;

[0031] Figure 6 It is a structural diagram of the electronic equipment provided by the application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical scheme and advantages of the application more clear, the technical scheme in the application will be described clearly and completely in combination with the drawings in the application. Obviously, the described embodiments are some embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the application.

[0033] It should be noted that in the description of the embodiments of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitation, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article or equipment comprising the element. The orientation or position relationship indicated by the terms "upper", "lower" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. Unless otherwise specified and limited, the terms "mounting", "connecting", "connecting" should be broadly understood, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be connected inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0034] In the deployment mode of the conventional neural network model, the self-constructed service needs the developer to write an engineering to load and hoist the model, and needs to be packaged after writing the input and output interfaces to complete the deployment, so that the model can be accessed and operated. It leads to long time consumption, high cost, and high requirement for professional accomplishment of deployment personnel.

[0035] The present application builds a model deployment platform, which can support automatic deployment of neural network models trained on parsed development frameworks. Users only need to understand the deployment requirements of the deployment platform and upload the deployment files of the neural network model to the model deployment platform. The model deployment platform can quickly obtain the calling ability of Restful API according to the model deployment method provided by the present application, thereby generating a web page address, an input interface and an output interface, and accelerating the deployment and verification process of the neural network model.

[0036] The following will be described in detail Figures 1 to 6 The model deployment method, identification method, device and electronic equipment provided by the embodiments of the present application are described.

[0037] Figure 1 The flowchart of the model deployment method provided by the present application is shown in Figure 1 applied to a model deployment platform (hereinafter referred to as: deployment platform), including but not limited to the following steps:

[0038] Firstly, in step S11, the deployment platform constructs a service image of the neural network model according to a deployment file of the neural network model.

[0039] In the case that the neural network model is an image recognition model, after the deployment platform predefines the operations on the test image, the neural network model and the input and output structure by using the configuration file, the configuration service in the deployment platform defines and sorts the order and parameters of the operations, decomposes the content, determines the specific called function, and concatenates the calling logic to obtain an inference instance, that is, an inference application container engine (docker) image.

[0040] The inference docker image can support inference on storage types of TensorFlow model, PyTorch model, Open Neural Network Exchange (ONNX) model and Compute Unified Device Architecture (CUDA) model. The storage type is related to the development platform.

[0041] Based on the inference docker image and the user's proprietary data set, a service docker image of the user is constructed.

[0042] The deployment file at least includes a model file of the neural network model, a configuration file and a test image.

[0043] The configuration file at least includes input information, output information and calling flow information of the neural network model. The input information includes the type and size requirement of the input image, such as adjusting the image size and image normalization processing. The inference instance of the deployment platform also includes the preprocessing of the to-be-recognized image to obtain a preprocessed image.

[0044] Since the image recognition model has mandatory requirements for the input to-be-recognized image, for example, the three-channel model requires the size of the to-be-recognized image to be 416*416, in order to facilitate the uploading of various images to test and call the image recognition model, the inference instance provides support for some image preprocessing functions, such as adjusting the size of the to-be-recognized image and normalizing the to-be-recognized image.

[0045] Figure 2 The model deployment platform provided by the application is shown in FIG. Figure 2 The model deployment platform provided by the application is shown in FIG.

[0046] The access layer at least includes a front-end interactive service, and the front-end interactive service provides an access entrance and an interaction interface for a user, and the user can perform functions such as login and logout of the system, user information management, data uploading, monitoring of a construction state, problem feedback and the like through the interface provided by the front-end interactive service. The data uploading includes uploading of a model file, a configuration file and a test file of a neural network model. In a case where the neural network model is an image recognition model, the test file can be a test image.

[0047] The interface provided by the front-end interactive service is also used to display a test result of the neural network model.

[0048] The business layer at least includes a work order management, a configuration service, a monitoring service, an interface service, a maintenance service and an inference instance.

[0049] The work order management at least includes service work order session management and maintenance work order session management. The service work order session management can perform compliance verification on a deployment file of a user to reject illegal access, and the maintenance work order session management can generate fault alarm information and a maintenance work order according to an error occurring in a running process of a deployment platform, so that a developer can perform specific operations or services according to an error type and process information in the maintenance work order by using the maintenance service.

[0050] The configuration service is used to organize and arrange a model file, a configuration file and a test image to generate an inference instance. The inference instance is responsible for providing online inference capability for a user, and can at least implement inference on a model, preprocessing on a to-be-recognized image and structural adjustment on output data of the model.

[0051] The monitoring service is used to monitor whether a process of constructing a service docker image by the configuration service is correct, to determine whether the constructed service docker image exists, and to determine whether the service docker image is normally started if the service docker image exists. If the service docker image fails in a construction process, failure information is returned in a timely manner.

[0052] The construction process of the service docker image is based on a container. Since the process of constructing the service docker image is only an operation of calling a container command, a state of the container in a calling process is not easy to monitor, and therefore the monitoring service is set to monitor the construction process of the service docker image.

[0053] The interface service is responsible for providing a unified interface call for a user, and is specifically used to define an input interface and an output interface of model deployment.

[0054] The maintenance service is used to perform a maintenance operation on a deployment platform.

[0055] The data layer is used for data storage and can at least include user data, work order data and logs. The work order data includes service work order data and operation and maintenance work order data.

[0056] K8s is used for deploying service images. K8s is a resource distributed management tool. For example, when a service needs to be started in the background, K8s management can dynamically scale the start according to the request amount to determine the number of services to start to achieve automatic management.

[0057] The deployment platform provides a dedicated access interface for users, and returns the result in the format of the image unified interface after inferring the data uploaded by the user. The deployment is a service docker image generated based on the basic inference docker image of the configuration service. The service docker image relies on K8s to provide deployment and load balancing, dynamic expansion.

[0058] For example, after obtaining the neural network model, the user wants to see the specific running effect of the model. For example, the user wants to verify whether a face recognition model accurately recognizes a face and the effect of running on hardware (such as a mobile phone or a chip). The model needs to be deployed, that is, the neural network model is converted into an input interface, an output interface and a web address. After the face recognition model is deployed, the recognition result output by the output interface can be whether there is a face in the to-be-recognized image. If the face in the to-be-recognized image is recognized, the recognition result also includes the specific position of the face.

[0059] Specifically, after the user fills in the configuration file of the neural network model, the model file, the configuration file and the test image are uploaded through the interface provided by the front-end interactive service. The business layer arranges and constructs a service image for the deployment file. After the service image is tested by the test image, whether the construction is successful is displayed on the interface provided by the front-end interactive service.

[0060] Further, in step S12, an inference instance is obtained by using a deployment tool of the model deployment platform according to the service image.

[0061] After the test image passes the test of the service docker image, the business layer of the deployment platform sends the service docker image to K8s for deployment.

[0062] K8s, as a deployment tool, is a service of a running image that is started. The business layer provides the service docker image and the corresponding instructions to K8s, and K8s can perform management operations on the service docker image, such as starting, expanding, stopping and the like, to generate an inference instance in the business layer.

[0063] Further, in step S13, the deployment platform generates a webpage link, an input interface and an output interface of the neural network model according to the inference instance.

[0064] The business layer generates a webpage link, an input interface and an output interface of the neural network model according to the inference instance generated by K8s.

[0065] The deployment of the model does not check whether the recognition result of the inference output is the output that the model really wants, for example: after the image test is completed, the entire image recognition process can be performed, and the model deployment is considered to be successful. Subsequently, the recognition result can be checked through the expansion of the platform.

[0066] The webpage address, the input interface and the output interface obtained by the deployment are used for testing, and after the test effect meets the standard, the neural network model can be placed on the hardware for development or calling.

[0067] The model deployment method provided by the application, the model deployment platform deploys the neural network model according to the deployment file provided by the user, generates corresponding webpage links, input interfaces and output interfaces, and the user only needs to understand the ability of the model deployment platform and learn to write the configuration file, so as to quickly deploy the model by using the model deployment platform, simplify the operation, improve the model deployment efficiency, reduce the threshold of model deployment, and be friendly to the user.

[0068] Optionally, the service image of the neural network model is constructed according to the deployment file of the neural network model, and specifically includes:

[0069] An inference image is generated according to the deployment file, and the deployment file includes a model file and a configuration file.

[0070] The deployment file is processed to generate a user-specific data set corresponding to the deployment file.

[0071] The service image of the neural network model is constructed according to the inference image and the specific data set.

[0072] The service image can be a service docker image. In the case that the neural network model is an image recognition model, the deployment file can further include a test file.

[0073] Specifically, after the business layer of the deployment platform predefines the operations on the test image, the neural network model and the input and output structure by using the configuration file, the configuration service in the business layer defines and sorts the order and parameters of the operations, decomposes the content, determines the specific called function, and connects the calling logic to obtain an inference instance, that is, an inference docker image.

[0074] The deployment file is organized and archived, a format supported by the deployment platform is generated, a user-specific data set is generated, and the specific data set is sent to the user data in the data layer for storage.

[0075] The configuration service of the deployment platform can construct a service docker image of the neural network model according to the inference docker image and the specific data set.

[0076] The model deployment method provided by the application generates an inference image through a deployment file, finally constructs a service image of a neural network model, simplifies human-computer interaction logic, and provides a basis for model deployment and rapid verification of model effects.

[0077] Optionally, the configuration file includes input information, output information, and calling process information of the neural network model.

[0078] When the neural network model is an image recognition model, the input information includes the name, size, and format of the input image of the neural network model.

[0079] The output information includes the data type of the recognition result output by the neural network model.

[0080] The calling process information includes the running logic of the neural network model.

[0081] Optionally, the model deployment method provided by the application further includes:

[0082] If the construction of the service image fails, an operation and maintenance work order is generated according to related information of the failure.

[0083] The operation and maintenance work order is sent to a background server.

[0084] Specifically, if a failure occurs in the running of the deployment platform, for example, a failure occurs in the construction process of the service image, the operation and maintenance work order session management in the deployment platform generates failure alarm information and an operation and maintenance work order according to related information of the failure. The operation and maintenance work order can include the time of failure generation, process information, and failure type.

[0085] The business layer of the deployment platform sends the failure alarm information to the interface provided by the front-end interactive service to report errors to the user, and sends the operation and maintenance work order to the background server. The developers use operation and maintenance services to perform specific operations or services according to the failure type and process information in the operation and maintenance work order, to repair the failure.

[0086] According to the model deployment method provided in the application, an operation and maintenance work order is generated from an error generated during the operation of the deployment platform, thereby providing a basis for repairing the error.

[0087] Optionally, the model deployment method provided in the application further comprises:

[0088] The compliance of the deployment file is checked, and if the deployment file is non-compliant, an illegal access prompt is generated.

[0089] Specifically, after a user uploads a deployment file on an interface provided by a front-end interactive service, a service work order session management in the deployment platform can check the compliance of the deployment file, and if the deployment file is non-compliant, access of the deployment file is rejected, and an illegal access prompt is generated and displayed on the interface provided by the front-end interactive service, so as to remind the user to adjust or replace the deployment file and re-upload.

[0090] According to the model deployment method provided in the application, the compliance of the deployment file is checked to limit illegal access, thereby ensuring the security of the deployment platform.

[0091] Figure 3 is a flowchart of the identification method provided in the application, as shown in Figure 3 comprises:

[0092] First, in step S31, a to-be-processed object input from an input interface is received, and the to-be-processed object is image data or voice data.

[0093] Further, in step S32, based on a web link, the to-be-processed object is identified by using an inference instance corresponding to a neural network model, to generate an identification result; the neural network model is used to identify the to-be-processed object.

[0094] Further, in step S33, the identification result is output by using an output interface; the inference instance, the web link, the input interface, and the output interface are determined based on the model deployment method as described above.

[0095] Specifically, a user can directly call a web address generated by the deployment platform, upload a to-be-identified image on an input interface, an interface service connects an inference instance, the inference instance feeds back an identification result of the to-be-identified image to the interface service, the interface service returns the result through an output interface, and the user can see the identification result of the neural network model on an interface provided by a front-end interactive service, and according to the test result, the accuracy of the model in identifying the test image is obtained, so as to evaluate and verify the neural network model.

[0096] According to the identification method provided by the application, the model is verified and evaluated by using the webpage link, the input interface and the output interface, so that the model developed by supporting multiple model platforms is supported, and load balancing and high concurrency are also supported.

[0097] Optionally, the identifying the to-be-processed object by using the inference instance specifically comprises:

[0098] The to-be-processed object is preprocessed by using the inference instance to obtain a preprocessed object.

[0099] The preprocessed object is identified to generate the identification result.

[0100] Specifically, if the neural network model is an image recognition model, the inference instance of the deployment platform further comprises preprocessing of a to-be-identified image to obtain a preprocessed image. Since the image recognition model has mandatory requirements for the input to-be-identified image, for example, a three-channel model requires that the size of the to-be-identified image is 416*416, in order to facilitate uploading of various images to test and call the image recognition model, the inference instance provides support for some image preprocessing functions, for example, adjusting the size of the to-be-identified image and normalizing the to-be-identified image.

[0101] Further, the inference instance of the deployment platform identifies the preprocessed image to generate an identification result.

[0102] According to the identification method provided by the application, the to-be-identified image is preprocessed by using the inference instance, so that the identification range is expanded, the processing data of the model is simplified, and the reliability of feature extraction, image segmentation, matching and identification is improved.

[0103] Figure 4 is a structural schematic diagram of a model deployment device provided by the application, as Figure 4 shown, comprising but not limited to:

[0104] The construction module 401 is configured to construct a service image of the neural network model according to a deployment file of the neural network model.

[0105] The acquisition module 402 is configured to acquire an inference instance by using a deployment tool of the model deployment platform according to the service image.

[0106] The first generation module 403 is configured to generate a webpage link, an input interface and an output interface of the neural network model according to the inference instance.

[0107] Firstly, the construction module 401 constructs a service image of the neural network model according to a deployment file of the neural network model.

[0108] The neural network model can be an image recognition model or a speech recognition model.

[0109] After the deployment platform predefines the operations on the test image, the neural network model and the input and output structure by using the configuration file, a configuration service in the deployment platform defines and sequences the order and parameters of the operations, decomposes the content, determines the specific called function, concatenates the calling logic, and obtains an inference instance, that is, an inference application container engine (docker) image.

[0110] The inference docker image can support inference on storage types of TensorFlow models, PyTorch models, ONNX models and CUDA models. The storage type is related to the development platform.

[0111] Based on the inference docker image and the user's private data set, a service docker image of the user is constructed.

[0112] The deployment file can include a model file of the neural network model, a configuration file and a test image.

[0113] The configuration file can include input information, output information and calling flow information of the neural network model. The input information includes the type and size requirement of the input image, such as adjusting the image size and image normalization processing.

[0114] The deployment platform includes an access layer, a business layer, a data layer and a deployment tool Kubernetes (K8s for short).

[0115] The access layer can include a front-end interactive service, which provides an access entry and an interactive interface for the user. The user can log in and log out of the system, manage user information, upload data, monitor the construction state, and feed back problems through the interface provided by the front-end interactive service. The data upload includes uploading the model file of the neural network model, the configuration file and the test image.

[0116] The interface provided by the front-end interactive service is also used to display the test result of the neural network model.

[0117] The business layer can include a work order management, a configuration service, a monitoring service, an interface service, an operation and maintenance service and an inference instance.

[0118] The work order management can at least include service work order session management and operation and maintenance work order session management. The service work order session management can perform compliance verification on the deployment file uploaded by the user to reject illegal access; and the operation and maintenance work order session management can generate fault alarm information and operation and maintenance work order according to errors occurring in the running process of the deployment platform, so that the developer can perform specific operation or service according to the error type and process information in the operation and maintenance work order by using the operation and maintenance service.

[0119] The configuration service is configured to organize and arrange the model file, the configuration file and the test image, and generate an inference instance. The inference instance is responsible for providing online inference capability for the user, and can at least realize inference on the model, pre-processing on the image to be recognized, and structure adjustment on the model output data.

[0120] The monitoring service is configured to monitor whether the process of building the service docker image by the configuration service is correct, and determine whether the built service docker image exists; if the built service docker image exists, it is determined whether the service docker image is normally started. If the service docker image fails in the building process, the failure information is returned in time.

[0121] The building process of the service docker image is based on a container. Since the process of building the service docker image is only an operation of calling a container command, the state of the container in the calling process is not easy to monitor, so the monitoring service is set to monitor the building process of the service docker image.

[0122] The interface service is responsible for providing a unified interface call for the user, and is specifically configured to define the input interface and the output interface of the model deployment.

[0123] The operation and maintenance service is configured to perform maintenance operation on the deployment platform.

[0124] The data layer is configured to store data, and can at least include user data, work order data and log.

[0125] The K8s is configured to deploy the service image. The K8s belongs to a resource distributed management tool. For example, in the case of starting a service in the background, the K8s management can dynamically scale the start according to the request amount, so as to determine the number of started services and realize automatic management.

[0126] The deployment platform provides a dedicated access interface for the user, and returns the result in the format of the image unified interface after performing inference on the data uploaded by the user. The deployment is a service docker image generated by the configuration service based on the basic inference docker image. The service docker image relies on the K8s to provide deployment and load balancing and dynamic expansion.

[0127] For example, after obtaining the neural network model, the user wants to see the specific running effect of the model, for example: the user wants to verify whether a face recognition model is accurate in recognizing faces, and the effect of running on hardware (such as a mobile phone or a chip), and needs to deploy the model, that is, to convert the neural network model into an input interface, an output interface, and a web address. After the face recognition model is deployed, the recognition result output by the output interface can be whether there is a face in the to-be-recognized image. If the face in the to-be-recognized image is recognized, the recognition result also includes the specific position of the face.

[0128] Specifically, after the user fills in the configuration file of the neural network model, the model file, the configuration file, and the test image are uploaded through the interface provided by the front-end interactive service. The business layer will organize and build a service image for the deployment file. After the service image is tested by the test image, whether the construction is successful is displayed on the interface provided by the front-end interactive service.

[0129] Further, the acquisition module 402 acquires an inference instance by using a deployment tool of the model deployment platform according to the service image.

[0130] After the test image passes the test on the service docker image, the business layer of the deployment platform sends the service docker image to K8s for deployment.

[0131] K8s, as a deployment tool, is a service of a running image that is started. The business layer provides the service docker image and the corresponding instructions to K8s, so that K8s can perform management operations on the service docker image, such as starting, expanding, and stopping, to generate an inference instance in the business layer.

[0132] Further, the first generation module 403 generates a web link, an input interface, and an output interface of the neural network model according to the inference instance.

[0133] The business layer generates a web link, an input interface, and an output interface of the neural network model according to the inference instance generated by K8s. The web link, the input interface, and the output interface are used to call the inference instance of the neural network model.

[0134] The deployment of the model does not verify whether the recognition result of the inference output is the output that the model really wants, for example: after the image is tested, the entire image recognition process can be performed, and the model deployment is considered to be successful. Subsequently, the verification of the recognition result can be realized by expanding the platform.

[0135] The web address, the input interface, and the output interface obtained by deployment are used for testing. After the test effect meets the standard, the neural network model can be developed or called on hardware.

[0136] The model deployment device provided by this invention deploys a neural network model based on a deployment file provided by the user, generating corresponding web page links, input interfaces, and output interfaces. Users only need to understand the capabilities of the model deployment platform and learn how to write configuration files to quickly deploy models using the platform. This simplifies operations, improves model deployment efficiency, lowers the barrier to entry for model deployment, and is user-friendly.

[0137] It should be noted that the model deployment apparatus provided in this embodiment of the invention can be implemented based on the model deployment method described in any of the above embodiments during specific execution, and this embodiment will not elaborate on this.

[0138] Figure 5 This is a schematic diagram of the structure of the identification device provided by the present invention, as shown below. Figure 5 As shown, it includes:

[0139] The receiving module 501 is used to receive the object to be processed input from the input interface, wherein the object to be processed is image data or voice data;

[0140] The second generation module 502 is used to identify the object to be processed based on the webpage link and the inference instance corresponding to the neural network model, and generate an identification result; the neural network model is used to identify the object to be processed.

[0141] Output module 503 is used to output the recognition result via an output interface;

[0142] The inference instance, the webpage link, the input interface, and the output interface are determined based on the model deployment method described above.

[0143] During device operation, receiving module 501 receives a user inputting a data to be processed object from the input interface. The data to be processed object is image data or voice data. Second generation module 502 identifies the data to be processed based on a webpage link and uses the inference instance corresponding to the neural network model to generate an identification result. The neural network model is used to identify the data to be processed. Output module 503 outputs the identification result using the output interface. The inference instance, the webpage link, the input interface, and the output interface are determined based on the model deployment method described above.

[0144] Specifically, if the neural network model is an image recognition model, the user can directly call the web address generated by the deployment platform, input the image to be recognized on the input interface, connect the inference instance through the interface service, and then feed back the recognition result of the image to be recognized to the interface service through the inference instance. The interface service returns the result through the output interface, so that the user can see the recognition result of the neural network model on the interface provided by the front-end interactive service, and obtain the accuracy of the model in recognizing the test image according to the test result, so as to evaluate and verify the neural network model.

[0145] According to the image recognition device provided by the application, the model verification and evaluation are realized by using the web link, the input interface and the output interface, which supports the models developed by various model platforms, and supports load balancing and high concurrency.

[0146] It should be noted that the image recognition device provided by the embodiment of the application can be realized based on the image recognition method of any of the above embodiments, and the embodiment will not be repeated.

[0147] Figure 6 is a structural schematic diagram of an electronic device provided by the application, as Figure 6 shown, the electronic device can include a processor 610, a communications interface 620, a memory 630 and a communications bus 640, wherein the processor 610, the communications interface 620 and the memory 630 communicate with each other through the communications bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the model deployment method, which includes: constructing a service image of the neural network model according to a deployment file of the neural network model; obtaining an inference instance by using a deployment tool of the model deployment platform according to the service image; and generating a web link, an input interface and an output interface of the neural network model according to the inference instance.

[0148] In addition, the logic instructions in the memory 630 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0149] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the model deployment method provided by the above-mentioned method, and the method comprises: constructing a service image of a neural network model according to a deployment file of the neural network model; obtaining an inference instance by using a deployment tool of the model deployment platform according to the service image; and generating a web link, an input interface and an output interface of the neural network model according to the inference instance.

[0150] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the model deployment method provided by the above-mentioned method, and the method comprises: constructing a service image of a neural network model according to a deployment file of the neural network model; obtaining an inference instance by using a deployment tool of the model deployment platform according to the service image; and generating a web link, an input interface and an output interface of the neural network model according to the inference instance.

[0151] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place, or distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0152] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0153] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model deployment method, applied to a model deployment platform, characterized in that, include: Based on the deployment file of the neural network model, constructing a service image of the neural network model includes: generating an inference image based on the deployment file, the deployment file including a model file and a configuration file; processing the deployment file to generate a user-specific dataset corresponding to the deployment file; constructing a service image based on the inference image and the specific dataset; and dynamically deploying the service image using Kubernetes, automatically adjusting the number of inference instances based on the request volume. Based on the service image, an inference instance is obtained using the deployment tool of the model deployment platform; the inference instance has an image preprocessing function, which includes adjusting the size of the image to be recognized and normalizing the image to be recognized; Based on the inference instance, generate the webpage link, input interface, and output interface of the neural network model; If the construction of the service image fails, an operation and maintenance work order is generated based on the relevant information of the failure; the operation and maintenance work order is then sent to the backend server. The deployment file is validated for compliance. If the deployment file is not compliant, an illegal access warning is generated.

2. The model deployment method according to claim 1, characterized in that, The configuration file includes the input information, output information, and calling process information of the neural network model; When the neural network model is an image recognition model, the input information includes the name, size, and format of the input image of the neural network model; The output information includes the data type of the recognition result output by the neural network model; The call process information includes the operating logic of the neural network model.

3. A recognition method, characterized in that, include: Receives an object to be processed from the input interface, wherein the object to be processed is image data or voice data; Based on a webpage link, the object to be processed is identified using the inference instance corresponding to the neural network model, and an identification result is generated. This includes: preprocessing the object to be processed using the inference instance to obtain a preprocessed object; identifying the preprocessed object using the inference instance to generate the identification result; and the neural network model is used to identify the object to be processed. The recognition result is output using the output interface; The inference instance, the webpage link, the input interface, and the output interface are determined based on the model deployment method as described in any one of claims 1 to 2.

4. A model deployment device, characterized in that, include: A build module is used to build a service image of the neural network model based on the deployment file of the neural network model, including: generating an inference image based on the deployment file, the deployment file including a model file and a configuration file; processing the deployment file to generate a user-specific dataset corresponding to the deployment file; building a service image based on the inference image and the specific dataset; and dynamically deploying the service image through Kubernetes, automatically adjusting the number of inference instances according to the request volume. The acquisition module is used to acquire an inference instance based on the service image and using the deployment tool of the model deployment platform; the inference instance has an image preprocessing function, which includes adjusting the size of the image to be recognized and normalizing the image to be recognized; The first generation module is used to generate the web page link, input interface and output interface of the neural network model based on the inference instance; If the construction of the service image fails, an operation and maintenance work order is generated based on the relevant information of the failure; the operation and maintenance work order is then sent to the backend server. The deployment file is validated for compliance. If the deployment file is not compliant, an illegal access warning is generated.

5. An identification device, characterized in that, include: The receiving module is used to receive the object to be processed input from the input interface, wherein the object to be processed is image data or voice data; The second generation module is used to identify the object to be processed based on the webpage link and the inference instance corresponding to the neural network model, and generate an identification result; the neural network model is used to identify the object to be processed. The output module is used to output the recognition result through the output interface; The inference instance, the webpage link, the input interface, and the output interface are determined based on the model deployment method as described in any one of claims 1 to 2.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the model deployment method as described in any one of claims 1 to 2, or the identification method as described in claim 3.

Citation Information

Patent Citations

  • Deployment and monitoring device and method for machine learning model

    CN111488254A

  • Inference service system based on Kubernetes

    CN111629061A

  • Model deployment method and device, computer equipment and storage medium

    CN112230911A