Model reasoning method and device based on deep learning, medium and electronic equipment

By converting the classification model into an accelerated inference model, combining the inference program and instantiation architecture, the problem of low model inference efficiency in the existing technology is solved, and efficient and resource-saving model inference effect is achieved.

CN120163232APending Publication Date: 2025-06-17JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311723151.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing deep learning-based model inference frameworks, such as using flask, have problems with low efficiency, which affects the performance of model inference.

Method used

Efficient model inference is achieved by converting the classification model into an accelerated inference model suitable for inference programs, and combining inference programs (e.g., triton) and instantiated architectures (e.g., tornado). This method does not require an inference program to preprocess the model. The instantiation architecture is used to process the inference results and output, which improves the overall model inference efficiency.

Benefits of technology

It realizes efficient model inference, reduces the consumption of equipment resources by the inference program, and has the advantages of convenient configuration and low development difficulty, improving the overall model inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163232A_ABST
    Figure CN120163232A_ABST
Patent Text Reader

Abstract

The invention provides a model reasoning method based on deep learning, a model reasoning device based on deep learning, a computer readable storage medium and electronic equipment, and relates to the technical field of machine learning. According to the method, the accelerated reasoning model can be conveniently and directly applied to the reasoning program, the reasoning program does not need to pre-process the model, efficient model reasoning can be achieved by combining the reasoning program and the instantiation framework, the reasoning program is only used for model reasoning, the instantiation framework is used for processing and outputting the reasoning result, the reasoning advantage of the reasoning program is exerted, and the reasoning efficiency is improved. Moreover, the model post-processing operation is realized through the instantiation architecture, so that the overall model reasoning efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology. Specifically, it relates to a model inference method based on deep learning, a model inference device based on deep learning, a computer-readable storage medium, and an electronic device. Background Art

[0002] Deep learning is a machine learning method based on artificial neural networks. It can automatically learn features and patterns from a large amount of data, thereby realizing various applications of artificial intelligence, such as image recognition, natural language processing, speech recognition, etc. The "depth" of deep learning refers to the number of layers of the neural network, usually having multiple hidden layers, and each layer can extract different levels of abstract representations of the data. The advantage of deep learning is that it does not require manual feature design, but automatically adjusts the parameters of the neural network through algorithms such as backpropagation and gradient descent, so that the model can gradually improve the accuracy of prediction or classification.

[0003] Currently, in the model inference service framework in the field of deep learning, a model is trained using a deep learning framework (such as tensorflow, pytorch), and then the inference model code is encapsulated in a way such as flask and provided for external calls in the form of an http service interface.

[0004] However, using inference frameworks such as flask will cause the problem of low efficiency in the inference process.

[0005] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of this application. Therefore, it may include information that does not constitute the existing solutions known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of this application is to provide a model inference method based on deep learning, a model inference device based on deep learning, a computer-readable storage medium, and an electronic device, which can convert a classification model into an accelerated inference model suitable for an inference program. Furthermore, it can facilitate the direct application of the accelerated inference model to the inference program without the inference program preprocessing the model. Combining an inference program (such as triton) and an instantiation architecture (such as tornado) can achieve efficient model inference. Among them, the inference program is only used for model inference, and the instantiation architecture is used to process the inference results and output them, giving play to the inference advantages of the inference program. In addition, through the instantiation architecture, post-processing operations of the model can be realized, which can improve the overall model inference efficiency.

[0007] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0008] According to one aspect of the present application, there is provided a model inference method based on deep learning, including:

[0009] Obtain a classification model based on deep learning;

[0010] Convert the classification model into an accelerated inference model suitable for an inference program;

[0011] Execute an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result;

[0012] Convert the inference result into a callable result through an instantiation architecture and output the callable result.

[0013] In an exemplary embodiment of the present application, obtaining a classification model based on deep learning includes:

[0014] Obtain a sample data set with labeled tags;

[0015] Construct a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier;

[0016] Train the neural network model based on the sample data set to obtain a classification model based on deep learning.

[0017] In an exemplary embodiment of the present application, converting the classification model into an accelerated inference model suitable for an inference program includes:

[0018] Convert the classification model into an exchange format model;

[0019] Map the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for the inference program.

[0020] In an exemplary embodiment of the present application, mapping the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for the inference program includes:

[0021] Perform model parsing on the exchange format model to obtain a parsing result;

[0022] Map the parameters in the parsing result into the corresponding network layers of the preset network architecture respectively to obtain an accelerated inference model suitable for the inference program; wherein, the versions of the inference program and the accelerated inference model are consistent.

[0023] In an exemplary embodiment of the present application, it further includes:

[0024] Test the accelerated inference model based on the test data packet corresponding to the preset network architecture.

[0025] In an exemplary embodiment of the present application, an inference process corresponding to an accelerated inference model is performed through an inference program to obtain an inference result, including:

[0026] Construct a model repository containing multi-level directories; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level;

[0027] Save the configuration file corresponding to the accelerated inference model in the model name level, and save the accelerated inference model in the model version number level;

[0028] Map the input of the accelerated inference model to the input tensor of the comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model;

[0029] Perform an inference process corresponding to the comprehensive model based on the input tensor, so that the comprehensive model generates an output tensor;

[0030] Determine the inference result corresponding to the accelerated inference model based on the output tensor.

[0031] In an exemplary embodiment of the present application, the inference result is converted into a callable result through an instantiation architecture, including:

[0032] Process the inference result and the label corresponding to the inference result into a callable result through the instantiation architecture; wherein, the callable result is called based on a specified protocol interface.

[0033] According to one aspect of the present application, there is provided a model inference device based on deep learning, including:

[0034] A classification model acquisition unit, configured to acquire a classification model based on deep learning;

[0035] A model conversion unit, configured to convert the classification model into an accelerated inference model applicable to the inference program;

[0036] A model inference unit, configured to perform an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result;

[0037] An inference result post-processing unit, configured to convert the inference result into a callable result through the instantiation architecture and output the callable result.

[0038] In an exemplary embodiment of the present application, the classification model acquisition unit acquires a classification model based on deep learning, including:

[0039] Acquire a sample data set with labeled tags;

[0040] Construct a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier;

[0041] Train a neural network model based on a sample data set to obtain a classification model based on deep learning.

[0042] In an exemplary embodiment of the present application, the model conversion unit converts the classification model into an accelerated inference model suitable for an inference program, including:

[0043] Convert the classification model into an exchange format model;

[0044] Map the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for an inference program.

[0045] In an exemplary embodiment of the present application, the model conversion unit maps the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for an inference program, including:

[0046] Perform model parsing on the exchange format model to obtain a parsing result;

[0047] Map the parameters in the parsing result into the corresponding network layers of the preset network architecture respectively to obtain an accelerated inference model suitable for an inference program; wherein, the versions of the inference program and the accelerated inference model are consistent.

[0048] In an exemplary embodiment of the present application, it further includes:

[0049] A test unit for testing the accelerated inference model based on a test data packet corresponding to the preset network architecture.

[0050] In an exemplary embodiment of the present application, the model inference unit executes an inference process corresponding to the accelerated inference model through an inference program to obtain an inference result, including:

[0051] Construct a model repository containing multi-level directories; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level;

[0052] Save the configuration file corresponding to the accelerated inference model in the model name level, and save the accelerated inference model in the model version number level;

[0053] Map the input of the accelerated inference model into the input tensor of a comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model;

[0054] Execute an inference process corresponding to the comprehensive model based on the input tensor to enable the comprehensive model to generate an output tensor;

[0055] Determine the inference result corresponding to the accelerated inference model based on the output tensor.

[0056] In an exemplary embodiment of the present application, the inference result post - processing unit converts the inference result into a callable result through an instantiated architecture, including:

[0057] Processing the inference result and the label corresponding to the inference result into a callable result through the instantiated architecture; wherein, the callable result is called based on a specified protocol interface.

[0058] According to one aspect of the present application, there is provided a computer - readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above - mentioned method is implemented.

[0059] According to one aspect of the present application, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the above - mentioned method by executing the executable instructions.

[0060] The exemplary embodiments of the present application may have some or all of the following beneficial effects:

[0061] In the deep - learning - based model inference method provided by an exemplary embodiment of the present application, a classification model can be converted into an accelerated inference model suitable for an inference program. Furthermore, it is convenient to directly apply the accelerated inference model to the inference program without the inference program pre - processing the model. Combining the inference program (such as, triton) and the instantiated architecture (such as, tornado) can achieve efficient model inference. Among them, the inference program is only used for model inference, and the instantiated architecture is used to process the inference result and output it, giving full play to the inference advantages of the inference program. In addition, by implementing the model post - processing operation through the instantiated architecture, the overall model inference efficiency can be improved. Moreover, since the present application can achieve efficient model inference by combining the inference program and the instantiated architecture, rather than only through the inference program to implement the overall model inference process, the consumption of device resources by the inference program can also be reduced. In addition, the method of combining the inference program and the instantiated architecture also has advantages such as convenient configuration and low development difficulty.

[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0064] Figure 1Schematically shows a flowchart of a deep learning-based model inference method according to an embodiment of the present application;

[0065] Figure 2 Schematically shows a schematic diagram of the classification model construction process according to an embodiment of the present application;

[0066] Figure 3 Schematically shows a schematic diagram of the accelerated inference model construction process according to an embodiment of the present application;

[0067] Figure 4 Schematically shows an input / output interface diagram of an onnx model according to an embodiment of the present application;

[0068] Figure 5 Schematically shows a schematic diagram of the application mode of the inference program combined with the instantiated architecture according to an embodiment of the present application;

[0069] Figure 6 Schematically shows a flowchart of a deep learning-based model inference method according to another embodiment of the present application;

[0070] Figure 7 Schematically shows a structural block diagram of a deep learning-based model inference device according to an embodiment of the present application;

[0071] Figure 8 Schematically shows a structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0072] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that one or more of the specific details can be omitted in practicing the technical solutions of this application, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of this application.

[0073] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0074] Please refer to Figure 1 , Figure 1 which schematically shows a flowchart of a deep learning-based model inference method according to an embodiment of the present application. As Figure 1 shown, the deep learning-based model inference method may include: step S110 to step S140.

[0075] Step S110: Obtain a deep learning-based classification model.

[0076] Step S120: Convert the classification model into an accelerated inference model suitable for an inference program.

[0077] Step S130: Execute an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result.

[0078] Step S140: Convert the inference result into a callable result through an instantiation architecture and output the callable result.

[0079] Implement Figure 1 the method shown, the classification model can be converted into an accelerated inference model suitable for an inference program. Furthermore, it is convenient to directly apply the accelerated inference model to the inference program without the inference program preprocessing the model. Combining the inference program (such as, triton) and the instantiation architecture (such as, tornado) can achieve efficient model inference. Among them, the inference program is only used for model inference, and the instantiation architecture is used to process the inference result and output it, giving play to the inference advantage of the inference program. In addition, through the instantiation architecture to implement the model post-processing operation, the overall model inference efficiency can be improved. In addition, since the present application can combine the inference program and the instantiation architecture to achieve efficient model inference, rather than only through the inference program to implement the overall model inference process, the consumption of device resources by the inference program can also be reduced. In addition, the method of combining the inference program and the instantiation architecture also has advantages such as convenient configuration and low development difficulty.

[0080] Next, the above functions of the present exemplary embodiment will be described in more detail.

[0081] In step S110, a deep learning-based classification model is obtained.

[0082] Specifically, a deep learning based classification model is a model that uses deep learning methods to solve problems such as text classification, image classification, and speech classification. It can classify data categories based on the features and labels of the input data. There are many types of deep learning based classification models, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), capsule networks, attention mechanisms, transformers, graph neural networks (GNNs), etc. In this application, the classification model can correspond to any classification objective, such as a text quality inspection classification objective. The specific architecture of the classification model can be configured according to actual needs, and the embodiments of this application do not make any limitations.

[0083] As an alternative embodiment, obtaining a deep learning based classification model includes: obtaining a sample data set with labeled tags; constructing a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier; and training the neural network model based on the sample data set to obtain a deep learning based classification model. This can achieve the training of a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier to determine a classification model that can achieve a specified classification objective.

[0084] Specifically, before obtaining the classification model, the sample data in the sample data set can be labeled according to preset tags (such as violations, warnings, compliance) to obtain a sample data set with labeled tags; among them, different sample data can correspond to different preset tags or the same preset tag, and the embodiments of this application do not make any limitations. When labeling the sample data, it can rely on a specified corpus annotation format, such as Question (sample data)\tlabel (preset tag). If Question is "There are still such big people, you know, they are already adults", the corresponding annotation result can be "There are still such big people, you know, they are already adults\twarning"; if Question is "Well, what's his illness? He's in the intensive care unit", the corresponding annotation result can be "Well, what's his illness? He's in the intensive care unit\tcompliance". In addition, this application does not limit the sample size of the sample data set (such as 7000).

[0085] Furthermore, after obtaining the sample data set with labeled tags, a neural network model (BertPretrain + finetune + softmax) can be constructed based on the pre-trained model (BertPretrain), the pre-trained parameter fine-tuning module (finetune), and the classifier (softmax). Input the sample data set with labeled tags into the neural network model and perform multiple rounds of training until convergence to obtain a classification model based on deep learning. Among them, the classification model based on deep learning can be saved as a model file in the.pt format.

[0086] Among them, BertPretrain is a class for pre-training the BERT model, and different data sets and parameters can be used to train the BERT model for downstream tasks. BertPretrain includes the masked language model pre-training task and the next sentence prediction pre-training task. Among them, the BERT model is an open-source pre-trained model that can use the architecture of the transformer to perform unsupervised learning pre-training on a large-scale unlabeled corpus. The goal of the BERT model is to train on a large-scale unlabeled corpus to obtain a text semantic representation (Representation) containing rich semantic information of the text, and then fine-tune the semantic representation of the text in a specific NLP task and finally apply it to this NLP task.

[0087] Among them, finetune is used to fine-tune the parameters of the model, that is, based on a pre-trained model, perform a small amount of training on it to adapt to a specific task or data set, and the general knowledge of the pre-trained model can be utilized to improve the performance of the model in a specific field. Based on the BERT pre-trained model, by adjusting some parameters of the model (BERTFinetune), it can adapt to specific tasks and data sets, which can utilize the language representation ability of BERT pre-training, save training time and computing resources, and at the same time obtain high task performance.

[0088] Among them, softmax is used to map a set of input values to a set of output values such that the output values are between [0, 1] and the sum of all output values is 1. Softmax is often used in multi-classification problems to convert the output of the model into a probability distribution.

[0089] Please refer to Figure 2 , Figure 2 which schematically shows a schematic diagram of the classification model construction process according to an embodiment of the present application. As Figure 2 shown, based on the pre-trained model (BertPretrain) 210, the pre-trained parameter fine-tuning module (finetune) 220, and the classifier (softmax) 230, the required classification model 240 can be trained.

[0090] In step S120, the classification model is converted into an accelerated inference model suitable for the inference program.

[0091] Specifically, the inference program can provide multi-framework machine learning services as an inference engine (e.g., triton), supporting models in various formats such as tensorflow, pytorch, and onnx. Triton provides high-performance, scalable, and manageable services, supporting two protocols, gRPC and REST, and functions such as model version control, dynamic loading and unloading, and multi-model concurrency. Among them, tensorflow is an open-source deep learning framework that can be used to build, train, and deploy various types of machine learning models. It provides support for multiple programming languages and platforms, as well as rich tools and libraries. Pytorch is another open-source deep learning framework that can also be used to build, train, and deploy various types of machine learning models, providing a flexible and easy-to-use programming interface. ONNX (Open Neural Network Exchange) is an open-source machine learning model exchange format that can be used to convert and share models between different frameworks. ONNX defines operators, data types, serialization methods, as well as some tools and libraries to support model conversion and optimization.

[0092] The accelerated inference model (TensorRT) suitable for the inference program is used to accelerate the inference of deep learning models. It can optimize the model into an efficient execution engine, improve performance, and reduce memory occupancy, supporting multiple deep learning frameworks such as tensorflow, pytorch, and onnx. Simply put, TensorRT can be a forward-propagation deep learning framework for accelerating and optimizing the trained model.

[0093] As an alternative embodiment, converting the classification model into an accelerated inference model suitable for the inference program includes: converting the classification model into an exchange format model; mapping the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for the inference program. This can convert the classification model into an accelerated inference model suitable for the inference program, so that the model inference results can be output more efficiently through the inference program, and the model inference efficiency can be improved.

[0094] Specifically, according to the TorchServe service system applied to the PyTorch model, the classification model (PyTorch model) can be converted into an interchange format model (ONNX model). Among them, TorchServer can provide high-performance, scalable, and manageable services, support two protocols, gRPC and REST, and support functions such as model version control, dynamic loading and unloading, and multi-model concurrency. Among them, the torch.onnx.export() function for model conversion is included in the service system of the PyTorch model.

[0095] As an alternative embodiment, mapping the interchange format model to a preset network architecture to obtain an accelerated inference model suitable for the inference program includes: parsing the interchange format model to obtain a parsing result; mapping the parameters in the parsing result to the corresponding network layers of the preset network architecture respectively to obtain an accelerated inference model suitable for the inference program; where the versions of the inference program and the accelerated inference model are consistent. In this way, efficient model conversion can be achieved through the provided model conversion method, and an accelerated inference model that can achieve accelerated inference can be obtained.

[0096] Specifically, TensorRT can parse the PyTorch model and map the parsed results one by one to the network layers in the preset network architecture (i.e., the TensorRT deep learning framework) to uniformly convert the interchange format model into TensorRT.

[0097] Please refer to Figure 3 , Figure 3 which schematically shows a schematic diagram of the accelerated inference model construction process according to an embodiment of the present application. As Figure 3 shown, the PyTorch model 310 can be first converted into an ONNX model 320, and then, the ONNX model 320 can be further converted into a TensorRT model 330, so that an accelerated inference model suitable for the inference program can be obtained.

[0098] Please refer to Figure 4 , Figure 4 which schematically shows an input / output interface diagram of the ONNX model according to an embodiment of the present application. Triggering the visualization tool (netrn) can display as Figure 4The interface diagram shown includes: the model information (ONNX**) converted in the format function, the information of the classification model (pytorch**) in the producder function, and the imported model file (ai.onnx**) in the imports function. Also, the input identification information (name:**) and data type (type:**) defined in the input identification field (input_ids), and the segment identification information (name:**) and data type (type:**) defined in the segment identification field (segment_ids). Furthermore, the output identification information (name:**) and data type (type:**) can be defined in the output field (output).

[0099] As an optional embodiment, it further includes: testing the accelerated inference model based on the test data packet corresponding to the preset network architecture. This can achieve the testing of the accelerated inference model to facilitate timely debugging of the accelerated inference model.

[0100] Specifically, according to the interface (api) of the preset network architecture (i.e., the tensorRT deep learning framework), the accelerated inference model implemented as a python package in tensorRT can be called to achieve the testing of the accelerated inference model. Furthermore, optionally, the accelerated inference model can be debugged based on the test results.

[0101] In step S130, the inference process corresponding to the accelerated inference model is executed through the inference program to obtain the inference result.

[0102] As an optional embodiment, the inference process corresponding to the accelerated inference model is executed through the inference program to obtain the inference result, including: constructing a model repository containing multi-level directories; where the multi-level directories at least include the repository name level, the model name level, and the model version number level; saving the configuration file corresponding to the accelerated inference model in the model name level, and saving the accelerated inference model in the model version number level; mapping the input of the accelerated inference model to the input tensor of the comprehensive model based on the model repository; where the comprehensive model is obtained by combining multiple models including the accelerated inference model; executing the inference process corresponding to the comprehensive model based on the input tensor to enable the comprehensive model to generate the output tensor; determining the inference result corresponding to the accelerated inference model based on the output tensor. This can improve the model inference efficiency.

[0103] Specifically, a model repository containing multi-level directories (e.g., three-level directories) can be constructed. The model repository can be used to record multiple models including the accelerated inference model (e.g., tensorflow model, pytorch model, onnx model, tensorRT model, etc.). Among them, each model corresponds to a configuration file, which is used to record the information and parameters required during the model inference process. Furthermore, the configuration file corresponding to the accelerated inference model can be saved in the model name level, and the accelerated inference model can be saved in the model version number level.

[0104] Specifically, the following code used to represent the model configuration can be referred to:

[0105]

[0106]

[0107] Specifically, the above code segment represents a model configuration file for a Triton server, which is used to define the attributes and parameters of a model named "xujia_engine".

[0108] Among them, the field "name" represents the model name, and here the model name is specified as "xujia_engine".

[0109] The field "platform" represents the backend type used by the model (e.g., TensorRT, TensorFlow, PyTorch, ONNX, etc.). Here, the backend type of the model is specified as "tensorrt_plan", indicating that the model is a TensorRT plan file.

[0110] "max_batch_size" represents the maximum batch size supported by the model. Here, max_batch_size: 0 indicates that the model does not support batch processing and can only process a single request.

[0111] The "input" represents a list of input tensors of the model. Each input tensor must specify a name, data type, and shape. Here, it is specified that the model has two input tensors, namely "input_ids" and "segment_ids". Among them, both "input_ids" and "segment_ids" are of integer type (TYPE_INT32), and both "input_ids" and "segment_ids" have two dimensions. The first dimension is variable (i.e., any positive integer) and is represented as -1, and the second dimension is fixed and represented as 128. Here, "ids" is an abbreviation for identifier, which is specifically used to distinguish the numbers of different text paragraphs. For example, the words in the first text can be represented by "ids" with a value of 0, and the words in the second text can be represented by "ids" with a value of 1.

[0112] The "output" represents a list of output tensors of the model. Each output tensor must specify a name, data type, and shape. Here, it is specified that the model has one output tensor "output", and "output" is of floating-point type (TYPE_FP32). Also, "output" has two dimensions. The first dimension is fixed and represented as 1, and the second dimension is fixed and represented as 4.

[0113] Furthermore, based on the model repository, the input of the accelerated inference model can be mapped to the input tensors of the comprehensive model (Pipeline model). Based on the input tensors, the inference process corresponding to the comprehensive model can be executed, so that the comprehensive model generates output tensors, and based on the output tensors, the inference result corresponding to the accelerated inference model can be determined. Among them, Ensemble can be understood as combining multiple models into one model to improve the accuracy and robustness of the model. The Triton server supports the Ensemble method and can support defining a model whose input and output are the input and output of other models. The method of connecting multiple models in a certain order to form a combined model can be called Pipelines. The Triton server supports Pipelines and can be used to define a Pipeline composed of multiple models, where each model can run on the Triton server. Specifically, the Triton server can implement Pipelines by using the Ensemble method, which can avoid the overhead of transmitting intermediate tensors, and the Ensemble method enables each model to run on the same server.

[0114] For the configuration of Ensemble, the following code can be referred to:

[0115]

[0116]

[0117] Specifically, the above code segment represents the configuration file of the Ensemble model of the Triton server, which is used to define the attributes and parameters of a Pipeline model composed of two models.

[0118] Among them, model_name represents the name of the customizable Ensemble model, "jingyin_preprocess_xujia_engine".

[0119] model_version represents the version number of the customizable Ensemble model. Here, the version number is specified as 1.

[0120] input_map represents the input mapping list of the Ensemble model, which is used to map the input tensors of the Ensemble model to the input tensors of its constituent models. Each input mapping must specify a key and a value. The key is the name of the input tensor of the constituent model, and the value is the name of the input tensor of the Ensemble model. Here, it is specified that the Ensemble model has one input mapping, that is, mapping the "input_0" tensor of the "jingyin_preprocess" model to the "TEXT" tensor of the Ensemble model.

[0121] output_map represents the output mapping list of the Ensemble model, which is used to map the output tensors of its constituent models to the output tensors of the Ensemble model. Each output mapping must specify a key and a value. The key is the name of the output tensor of the Ensemble model, and the value is the name of the output tensor of the constituent model. Here, it is specified that the Ensemble model has three output mappings, which map the "input_ids" and "segment_ids" tensors of the "jingyin_preprocess" model to the "PREPROC_EMBEDDING" and "PREPROC_EMBEDDING_1" tensors of the Ensemble model respectively, and map the "output" tensor of the "xujia_engine" model to the "CLASSIFICATION" tensor of the Ensemble model.

[0122] In step S140, the inference result is converted into a callable result through instantiating the architecture and the callable result is output.

[0123] Specifically, the instantiated architecture can be used as a service construction framework (e.g., Tornado), a Python framework for building asynchronous web applications, which can handle a large number of concurrent connections and provide high-performance network I / O and web services. Specifically, by using non-blocking network input / output (I / O), Tornado can scale a large number of open connections, making it suitable for long polling, Socket, and other applications that require long-term connections with each user. In this application, the instantiated architecture (Tornado) can be specifically used to implement functions such as front-end data preprocessing and back-end label conversion for the model.

[0124] As an alternative embodiment, converting the inference result into a callable result through the instantiated architecture includes: processing the inference result and the label corresponding to the inference result into a callable result through the instantiated architecture; wherein, the callable result is obtained based on a specified protocol interface. This can facilitate users to obtain the callable result and improve the convenience of calling.

[0125] Specifically, after the inference program generates the inference result, the instantiated architecture can test and call the inference result through the Http interface / gRPC interface. Furthermore, the instantiated architecture can process the inference result and the label corresponding to the inference result into a callable result and output it based on the logic defined by the specified rules; wherein, the specified rules can be set arbitrarily according to actual needs. Furthermore, the callable result to be output can be called through the specified protocol interface, and the specified protocol interface can be set according to actual needs.

[0126] Please refer to Figure 5 , Figure 5 which schematically shows a schematic diagram of the application mode of the inference program combined with the instantiated architecture according to an embodiment of the present application.

[0127] As Figure 5 shown, a model repository (Models) 510 containing multi-level directories can be constructed first; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level; furthermore, the configuration file corresponding to the accelerated inference model is saved in the model name level, and the accelerated inference model is saved in the model version number level. Furthermore, the input of the accelerated inference model Model1 is mapped to the input tensor of the comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model.

[0128] The inference program (triton) 520 can be used to infer an integrated model based on TensorRT 530. The integrated model includes Model1, so that the integrated model generates an output tensor and determines an inference result corresponding to the accelerated inference model based on the output tensor. The inference process is specifically based on the configuration file Model1-engine:config.pbtxt 531, the post-processing logic Model1-backend 532, and the integrated model information Model1-ensemble:config.pbtxt 533. Here, backend refers to writing and calling the functions and methods of the model using an abstract interface.

[0129] Furthermore, a test call can be provided to the instantiated architecture (tornado) 550 based on the Http interface / gRPC interface 540. The instantiated architecture (tornado) 550 can execute a judgment logic 560 based on the post-processing logic Model1-backend 532, that is, judge whether there is a rule. If so, execute the backend logic processing 570 based on the post-processing logic Model1-backend 532 to process the inference result 580 and the label corresponding to the inference result 580 into a callable result and output it; if not, directly output the inference result 580.

[0130] Please refer to Figure 6 , Figure 6 which schematically shows a flowchart of a deep learning-based model inference method according to another embodiment of the present application. As Figure 6 shown, the deep learning-based model inference method may include: step S600 to step S624.

[0131] Step S600: Obtain a sample data set with labeled tags.

[0132] Step S602: Build a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier.

[0133] Step S604: Train the neural network model based on the sample data set to obtain a deep learning-based classification model.

[0134] Step S606: Convert the classification model into an exchange format model.

[0135] Step S608: Parse the exchange format model to obtain a parsing result.

[0136] Step S610: Map the parameters in the parsing result to the corresponding network layers of a preset network architecture respectively to obtain an accelerated inference model applicable to the inference program; where the versions of the inference program and the accelerated inference model are consistent.

[0137] Step S612: Test the accelerated inference model based on the test data packet corresponding to the preset network architecture.

[0138] Step S614: Build a model repository containing multi-level directories; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level.

[0139] Step S616: Save the configuration file corresponding to the accelerated inference model in the model name level, and save the accelerated inference model in the model version number level.

[0140] Step S618: Map the input of the accelerated inference model to the input tensor of the comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model.

[0141] Step S620: Execute the inference process corresponding to the comprehensive model based on the input tensor, so that the comprehensive model generates an output tensor.

[0142] Step S622: Determine the inference result corresponding to the accelerated inference model based on the output tensor.

[0143] Step S624: Process the inference result and the label corresponding to the inference result into a callable result through an instantiated architecture and output the callable result; wherein, the callable result is called based on a specified protocol interface.

[0144] It should be noted that steps S600 to S624 correspond to Figure 1 the steps and their embodiments shown. For the specific implementation manners of steps S600 to S624, please refer to Figure 1 the steps and their embodiments shown, which will not be elaborated here.

[0145] It can be seen that implementing Figure 6The method shown can convert a classification model into an accelerated inference model suitable for an inference program. Furthermore, it can facilitate the direct application of the accelerated inference model to the inference program without the need for the inference program to preprocess the model. Combining the inference program (e.g., triton) and the instantiated architecture (e.g., tornado) can achieve efficient model inference. Among them, the inference program is only used for model inference, and the instantiated architecture is used to process the inference results and output them, giving play to the inference advantages of the inference program. In addition, by implementing post-processing operations on the model through the instantiated architecture, the overall model inference efficiency can be improved. Moreover, since this application can achieve efficient model inference by combining the inference program and the instantiated architecture, rather than only through the inference program to implement the overall model inference process, it can also reduce the consumption of device resources by the inference program. In addition, the method of combining the inference program and the instantiated architecture also has advantages such as convenient configuration and low development difficulty.

[0146] Please refer to Figure 7 , Figure 7 which schematically shows a structural block diagram of a deep learning-based model inference device according to an embodiment of the present application. As Figure 7 shown, the deep learning-based model inference device 700 includes: a classification model acquisition unit 710, a model conversion unit 720, a model inference unit 730, and an inference result post-processing unit 740.

[0147] The classification model acquisition unit 710 is used to acquire a deep learning-based classification model;

[0148] The model conversion unit 720 is used to convert the classification model into an accelerated inference model suitable for the inference program;

[0149] The model inference unit 730 is used to execute the inference process corresponding to the accelerated inference model through the inference program to obtain an inference result;

[0150] The inference result post-processing unit 740 is used to convert the inference result into a callable result and output the callable result through the instantiated architecture.

[0151] It can be seen that implementing Figure 7The device shown can convert a classification model into an accelerated inference model suitable for an inference program. Furthermore, it can facilitate the direct application of the accelerated inference model to the inference program without the need for the inference program to preprocess the model. Combining the inference program (e.g., triton) and the instantiated architecture (e.g., tornado) can achieve efficient model inference. Among them, the inference program is only used for model inference, and the instantiated architecture is used to process the inference results and output them, giving play to the inference advantages of the inference program. In addition, by implementing the model post-processing operation through the instantiated architecture, the overall model inference efficiency can be improved. Moreover, since this application can achieve efficient model inference by combining the inference program and the instantiated architecture, rather than only through the inference program to implement the overall model inference process, the consumption of device resources by the inference program can also be reduced. In addition, the method of combining the inference program and the instantiated architecture also has advantages such as convenient configuration and low development difficulty.

[0152] In an exemplary embodiment of the present application, the classification model acquisition unit 710 acquires a classification model based on deep learning, including:

[0153] Acquire a sample data set with labeled tags;

[0154] Construct a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier;

[0155] Train the neural network model based on the sample data set to obtain a classification model based on deep learning.

[0156] It can be seen that implementing this optional embodiment can achieve the training of a neural network model constructed based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier to determine a classification model that can achieve the specified classification target.

[0157] In an exemplary embodiment of the present application, the model conversion unit 720 converts the classification model into an accelerated inference model suitable for the inference program, including:

[0158] Convert the classification model into an exchange format model;

[0159] Map the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for the inference program.

[0160] It can be seen that implementing this optional embodiment can convert the classification model into an accelerated inference model suitable for the inference program, so that the model inference result can be output more efficiently through the inference program, and the model inference efficiency can be improved.

[0161] In an exemplary embodiment of the present application, the model conversion unit 720 maps the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for the inference program, including:

[0162] Parse the exchange format model to obtain a parsing result;

[0163] Map the parameters in the parsing result to the corresponding network layers of a preset network architecture respectively to obtain an accelerated inference model applicable to an inference program; wherein, the versions of the inference program and the accelerated inference model are consistent.

[0164] It can be seen that implementing this optional embodiment can achieve efficient model conversion through the provided model conversion method, and obtain an accelerated inference model that can achieve accelerated inference.

[0165] In an exemplary embodiment of the present application, it further includes:

[0166] A test unit for testing the accelerated inference model based on a test data packet corresponding to the preset network architecture.

[0167] It can be seen that implementing this optional embodiment can achieve the test of the accelerated inference model, so as to facilitate timely debugging of the accelerated inference model.

[0168] In an exemplary embodiment of the present application, the model inference unit 730 executes an inference process corresponding to the accelerated inference model through an inference program to obtain an inference result, including:

[0169] Construct a model repository containing multi-level directories; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level;

[0170] Save the configuration file corresponding to the accelerated inference model in the model name level, and save the accelerated inference model in the model version number level;

[0171] Map the input of the accelerated inference model to the input tensor of a comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model;

[0172] Execute an inference process corresponding to the comprehensive model based on the input tensor to enable the comprehensive model to generate an output tensor;

[0173] Determine the inference result corresponding to the accelerated inference model based on the output tensor.

[0174] It can be seen that implementing this optional embodiment can improve the model inference efficiency.

[0175] In an exemplary embodiment of the present application, the inference result post-processing unit 740 converts the inference result into a callable result through an instantiated architecture, including:

[0176] The inference result and the label corresponding to the inference result are processed into a callable result through an instantiated architecture; wherein, the callable result is obtained for calling based on a specified protocol interface.

[0177] It can be seen that implementing this optional embodiment can facilitate the user to obtain the callable result and improve the calling convenience.

[0178] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0179] Since each functional module of the deep learning-based model inference device in the exemplary embodiments of the present application corresponds to the functions of the exemplary embodiments of the above-mentioned deep learning-based model inference method, for details not disclosed in the device embodiments of the present application, please refer to the embodiments of the above-mentioned deep learning-based model inference method of the present application.

[0180] Please refer to Figure 8 , Figure 8 which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.

[0181] It should be noted that Figure 8 the computer system 800 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0182] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0183] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as required so that a computer program read therefrom is installed into the storage section 808 as required.

[0184] Specifically, according to an embodiment of the present application, the processes described in the above reference flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, various functions defined in the methods and apparatuses of the present application are executed.

[0185] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the methods described in the above embodiments.

[0186] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0187] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0188] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0189] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include known common knowledge or conventional technical means in the art not disclosed in this application. The specification and examples are only to be considered as exemplary, and the true scope and spirit of this application are pointed out by the foregoing claims.

Claims

1. A method for model inference based on deep learning, characterized in that, including: Obtain a classification model based on deep learning; Convert the classification model into an accelerated inference model suitable for an inference program; Execute an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result; Convert the inference result into a callable result through an instantiated architecture and output the callable result.

2. The method according to claim 1, characterized in that, Obtain a classification model based on deep learning, including: Obtain a sample data set with labeled tags; Construct a neural network model based on a pre-trained model, a pre-trained parameter fine-tuning module, and a classifier; Train the neural network model based on the sample data set to obtain a classification model based on deep learning.

3. The method according to claim 1, characterized in that, Convert the classification model into an accelerated inference model suitable for an inference program, including: Convert the classification model into an exchange format model; Map the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for an inference program.

4. The method according to claim 3, characterized in that, Map the exchange format model into a preset network architecture to obtain an accelerated inference model suitable for an inference program, including: Perform model parsing on the exchange format model to obtain a parsing result; Map the parameters in the parsing result to the corresponding network layers of the preset network architecture respectively to obtain an accelerated inference model suitable for an inference program; wherein, the versions of the inference program and the accelerated inference model correspond.

5. The method according to claim 3, characterized in that, Also included: Test the accelerated inference model based on the test data packet corresponding to the preset network architecture.

6. The method according to claim 1, characterized in that, Execute an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result, including: Construct a model repository containing multi-level directories; wherein, the multi-level directories at least include a repository name level, a model name level, and a model version number level; Save the configuration file corresponding to the accelerated inference model in the model name level, and save the accelerated inference model in the model version number level; Map the input of the accelerated inference model into the input tensor of a comprehensive model based on the model repository; wherein, the comprehensive model is obtained by combining multiple models including the accelerated inference model; Execute an inference process corresponding to the comprehensive model based on the input tensor to enable the comprehensive model to generate an output tensor; Determine the inference result corresponding to the accelerated inference model based on the output tensor.

7. The method according to any one of claims 1 to 6, characterized in that, Convert the inference result into a callable result through an instantiated architecture, including: Convert the inference result and the label corresponding to the inference result into a callable result through an instantiated architecture; wherein, the callable result is called based on a specified protocol interface.

8. A model inference device based on deep learning, characterized in that, including: A classification model acquisition unit for obtaining a classification model based on deep learning; A model conversion unit for converting the classification model into an accelerated inference model suitable for an inference program; A model inference unit for executing an inference process corresponding to the accelerated inference model through the inference program to obtain an inference result; An inference result post-processing unit for converting the inference result into a callable result through an instantiated architecture and outputting the callable result.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that,When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method according to any one of claims 1 to 7 by executing the executable instructions.