Method for Operating Inference Model, Back End, Storage Medium, and Electronic Device

By segmenting the inference model into multiple sub-inference models and deploying it on multiple nodes, the inference model deployment failure caused by insufficient single node resources is solved, and the utilization rate of graphics processor resources is improved.

CN116822630BActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310787918.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-05-30
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

In the prior art, a single node cannot meet the graph processor resource requirements of the inference model, resulting in the failure of the inference model deployment.

Method used

The inference model is segmented into multiple sub-inference models. Each sub-inference model has a small resource requirement and is deployed on multiple nodes that meet the resource requirements. Each sub-inference model is run in a preset order through multiple second pods to complete the inference process of the entire inference model.

Benefits of technology

This avoids the failure of inference model deployment caused by insufficient single-node graphics processor resources, and improves the utilization rate of graphics processor resources distributed in multiple nodes in the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116822630B_ABST
    Figure CN116822630B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method for running an inference model, a backend, a storage medium, and an electronic device. The method includes: when a first Pod receives segmentation information and an inference model from a front end, the first Pod divides the inference model into multiple sub-inference models according to the segmentation information; the first Pod schedules a second Pod to a corresponding node and deploys the sub-inference models in the corresponding second Pods; the first Pod sends the input data of the inference model to the second Pod running the first sub-inference model and obtains the output data of the inference model from the second Pod running the last sub-inference model. Through the present application, the problem that a single node in the prior art cannot meet the graphics processor resource requirements of the inference model, resulting in the failure of the inference model deployment, is solved, and the utilization rate of the graphics processor resources distributed on multiple nodes in the cluster is improved accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computers, and more specifically, to a method for running an inference model, a method for running an inference model, a backend, a computer-readable storage medium, and an electronic device. Background Art

[0002] Currently, the steps for deploying an inference model on an inference platform are as follows: 1. The user uploads the inference model on the inference platform, and after the inference model is uploaded, it is saved in persistent storage; 2. The user uses the address of the uploaded inference model on the interface for creating a service on the inference platform, fills in parameters such as the required graphics processor resources, and then submits a request for creating a service; 3. After receiving the request for creating a service, the inference platform deploys the inference model according to the parameters such as the inference model and the graphics processor resources in the request for creating a service. When the inference platform deploys the inference model, it uses a Kubernetes Pod to run the inference model. When Kubernetes schedules the Pod, it checks the graphics processor resource information of all nodes in the cluster, finds a single node that can meet the graphics processor resources required by the model, and then schedules the Pod of the service to the node. 4. After the inference model is deployed as a service, real-time data can use this service for inference prediction.

[0003] When the graphics processor resources included in a node meet the inference model, the service can be successfully scheduled, but when the graphics processor resources of a node cannot meet the inference service, the service cannot be successfully scheduled.

[0004] As can be seen from the above, in the prior art, there is a problem that the inference model deployment fails because a single node cannot meet the graphics processor resource requirements of the inference model. Summary of the Invention

[0005] Embodiments of the present application provide a method for running an inference model, a method for running an inference model, a backend, a computer-readable storage medium, and an electronic device, so as to at least solve the problem that the inference model deployment fails because a single node cannot meet the graphics processor resource requirements of the inference model in the prior art.

[0006] According to an embodiment of the present application, a method for running an inference model is provided. A first Pod and multiple second Pods are configured in the backend. The first Pod is communicatively connected to the front end. The method is applied to the first Pod. The inference model includes multiple neural network layers, and the neural network layers are connected in sequence. The method includes: when the first Pod receives the segmentation information from the front end and the inference model, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information. The segmentation information corresponds to the sub-inference models one by one. One sub-inference model includes multiple neural network layers. One piece of segmentation information includes the name of the first neural network layer of the corresponding sub-inference model and the name of the last neural network layer of the corresponding sub-inference model; the first Pod schedules the second Pods to the corresponding nodes and deploys the sub-inference models in the corresponding second Pods. One node is configured with at least one graphics processor. The second Pods correspond to the nodes one by one, and the second Pods correspond to the sub-inference models one by one. The second Pod is used to run the corresponding sub-inference model and send the output data of the corresponding sub-inference model to the second Pod that runs the next sub-inference model. The multiple sub-inference models run in a preset order; the first Pod sends the input data of the inference model to the second Pod that runs the first sub-inference model and obtains the output data of the inference model from the second Pod that runs the last sub-inference model.

[0007] In an exemplary embodiment, there are M pieces of segmentation information. When the first Pod receives the segmentation information from the front end and the inference model, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information, including: the first Pod executes an extraction step. According to the Nth piece of segmentation information, the Nth sub-inference model is extracted from the inference model and the Nth sub-inference model is added to the scheduling queue, where 1 ≤ N ≤ M and N is a positive integer. The Nth piece of segmentation information includes the name of the first neural network layer of the Nth sub-inference model and the name of the last neural network layer of the Nth sub-inference model; the first Pod repeats the extraction step at least once until M sub-inference models are added to the scheduling queue. The output of the previous sub-inference model in the scheduling queue is the input of the next sub-inference model.

[0008] In an exemplary embodiment, the inference model includes M sub-inference models. The first Pod schedules the second Pod to a corresponding node and deploys the sub-inference model in the corresponding second Pod, including: The first Pod executes a creation step to create the Nth second Pod, where 1 ≤ N ≤ M and N is a positive integer; The first Pod executes a processing step to determine the Nth node according to a first quantity, and schedules the Nth second Pod to the Nth node. The first quantity is the number of graphics processors required to run the Nth sub-inference model, and the first quantity is less than or equal to a second quantity, where the second quantity is the number of graphics processors of the Nth node; The first Pod executes a first acquisition step to obtain the Nth sub-inference model from a scheduling queue. According to the Nth sub-inference model, determine the model weight of the Nth sub-inference model and the network structure of the Nth sub-inference model. The output of the previous sub-inference model in the scheduling queue is the input of the next sub-inference model; The first Pod executes a first sending step to send the model weight of the Nth sub-inference model, the network structure of the Nth sub-inference model, and the IP address of the (N + 1)th sub-inference model to the Nth second Pod; The first Pod executes a start step to start the Nth inference model; The first Pod repeats the execution of the creation step, the processing step, the first acquisition step, the first sending step, and the start step at least once until the start of M inference models is completed.

[0009] In an exemplary embodiment, after the first Pod schedules the second Pod to a corresponding node and deploys the sub-inference model in the corresponding second Pod, the method further includes: The first Pod sends a plurality of first data packets to the second Pod running the first sub-inference model, and obtains a plurality of second data packets from the second Pod running the last sub-inference model. The first data packet includes the input data and identifier of the inference model, and the second data packet includes the output data and identifier of the inference model. One first data packet corresponds to one input data of the inference model, one second data packet corresponds to one output data of the inference model, one input data of the inference model corresponds to one identifier, and one input data of the inference model corresponds to one output data of the inference model.

[0010] According to another embodiment of the present application, a method for running an inference model is provided. The inference model is divided into M sub-inference models, and a first Pod and multiple second Pods are configured in the backend. The Nth second Pod among the multiple second Pods is the second Pod that runs the Nth sub-inference model. The sub-inference models run in a preset order, where N is a positive integer natural number. The method is applied to the multiple second Pods, and the method includes: The first second Pod among the multiple second Pods receives the input data of the inference model from the first Pod; The first second Pod among the multiple second Pods inputs the input data of the inference model into the first sub-inference model to obtain the output data of the first sub-inference model; The first second Pod among the multiple second Pods sends the output data of the first sub-inference model to the second second Pod according to the IP address of the second second Pod; A first receiving step, the Kth second Pod among the multiple second Pods receives the output data of the (K - 1)th sub-inference model from the (K - 1)th second Pod, where 2 ≤ K ≤ M and K is a positive integer natural number; A first control step, the Kth second Pod among the multiple second Pods inputs the output data of the (K - 1)th sub-inference model into the Kth sub-inference model to obtain the output data of the Kth sub-inference model; A second sending step, the Kth second Pod among the multiple second Pods sends the output data of the Kth sub-inference model to the (K + 1)th second Pod according to the IP address of the (K + 1)th second Pod; A first determining step, the Kth second Pod among the multiple second Pods determines whether the IP address of the (K + 1)th second Pod is the IP address of the first Pod; Repeat the first receiving step, the first control step, the second sending step, and the first determining step at least once until the IP address of the (K + 1)th second Pod is the IP address of the first Pod and ends.

[0011] In an exemplary embodiment, the method further includes: the first second Pod among the multiple second Pods receives a third data packet from the first Pod, where the third data packet includes: input data of the inference model and an identifier; the first second Pod among the multiple second Pods inputs the input data of the inference model into the first sub-inference model to obtain output data of the first sub-inference model; the first second Pod among the multiple second Pods packs the output data of the first sub-inference model and the identifier into a fourth data packet; the first second Pod among the multiple second Pods sends the fourth data packet to the second second Pod according to the address of the second second Pod; a second receiving step, the K-th second Pod among the multiple second Pods receives a fifth data packet from the (K - 1)-th second Pod, where the fifth data packet includes output data of the (K - 1)-th sub-inference model and the identifier; a second control step, the K-th second Pod among the multiple second Pods inputs the output data of the (K - 1)-th sub-inference model into the K-th sub-inference model to obtain output data of the K-th sub-inference model; a packing step, the K-th second Pod among the multiple second Pods packs the output data of the K-th sub-inference model and the identifier into a sixth data packet; a third sending step, the K-th second Pod among the multiple second Pods sends the sixth data packet to the (K + 1)-th second Pod according to the address of the (K + 1)-th second Pod; a second determination step, the K-th second Pod among the multiple second Pods determines whether the IP address of the (K + 1)-th second Pod is the IP address of the first Pod; repeat the second receiving step, the second control step, the packing step, the third sending step, and the second determination step at least once until the IP address of the (K + 1)-th second Pod is the IP address of the first Pod to end.

[0012] In an exemplary embodiment, the K-th second Pod among the multiple second Pods packs the output data of the K-th sub-inference model and the identifier into a sixth data packet, including: the K-th second Pod among the multiple second Pods compresses the output data of the K-th sub-inference model, and packs the compressed output data of the K-th sub-inference model and the identifier into the sixth data packet.

[0013] According to another embodiment of the present application, a backend is provided. The inference model includes multiple neural network layers, and the neural network layers are connected in sequence. The backend includes: a first Pod, which is communicatively connected to the frontend; a second Pod; wherein, when the first Pod receives the segmentation information from the frontend and the inference model, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information. The segmentation information corresponds to the sub-inference models one by one. One sub-inference model includes multiple neural network layers. One piece of segmentation information includes the name of the first neural network layer of the corresponding sub-inference model and the name of the last neural network layer of the corresponding sub-inference model. The first Pod schedules the second Pod to the corresponding node and deploys the sub-inference model on the corresponding second Pod. One node is configured with at least one graphics processor. The second Pod corresponds to the node one by one, and the second Pod corresponds to the sub-inference model one by one. The second Pod is used to run the corresponding sub-inference model and send the output data of the corresponding sub-inference model to the second Pod that runs the next sub-inference model. The multiple sub-inference models run in a preset order. The first Pod sends the input data of the inference model to the second Pod that runs the first sub-inference model and obtains the output data of the inference model from the second Pod that runs the last sub-inference model.

[0014] According to still another embodiment of the present application, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0015] According to an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0016] Through the present application, the inference model is divided into multiple sub-inference models. The graphics processor resource requirements of each sub-inference model are small. Each sub-inference model can find a node that meets the graphics processor resource requirements of the sub-inference model. Each sub-inference model runs a part of the inference, and all sub-inference models jointly complete the inference process of the entire inference model. Furthermore, when the graphics processor resources included in one node cannot meet the requirements of the inference model, the situation where the inference model deployment fails is avoided, and the problem in the prior art that the inference model deployment fails because the graphics processor resources of a single node cannot meet the requirements of the inference model is solved, achieving the effect of improving the utilization rate of the graphics processor resources distributed on multiple nodes in the cluster. Description of the Drawings

[0017] Figure 1 is a flowchart of a method for running an inference model according to an embodiment of the present application;

[0018] Figure 2 is a schematic diagram of the segmentation of an inference model according to an embodiment of the present application;

[0019] Figure 3 is a flowchart of the deployment of an inference model according to an embodiment of the present application;

[0020] Figure 4 is a flowchart of a method for running another inference model according to an embodiment of the present application;

[0021] Figure 5 is a flowchart of the operation of an inference model according to an embodiment of the present application. Detailed Embodiments

[0022] Embodiments of the present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0024] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:

[0025] Kubernetes: An open-source container orchestration platform that can automate the deployment, scaling, and shrinking of containers, as well as provide more flexible environment configuration and automated management.

[0026] Pod: In Kubernetes, the smallest deployable unit is a Pod, which is a collection of one or more containers that share the same network namespace and storage volume. Usually, all containers running in the same Pod belong to the same application and need to cooperate to complete the tasks of the application.

[0027] Cluster: A cluster consists of multiple nodes.

[0028] Runtime: Refers to all the code libraries, frameworks, platforms, etc. required when the sub-inference model runs.

[0029] The first Pod and multiple second Pods are configured in the backend, and the first Pod is communicatively connected to the frontend.

[0030] In this embodiment, a method of running on the first Pod is provided. The inference model includes multiple neural network layers, and the neural network layers are connected in sequence. Figure 1 It is a flowchart according to an embodiment of the present application, as Figure 1 shown. The process includes the following steps:

[0031] Step S101, when the first Pod receives the segmentation information from the front end and the inference model, the first Pod divides the inference model into multiple sub-inference models according to each segmentation information.

[0032] Among them, the segmentation information corresponds to the sub-inference models one by one. One sub-inference model includes multiple neural network layers. One piece of segmentation information includes the name of the first neural network layer of the corresponding sub-inference model and the name of the last neural network layer of the corresponding sub-inference model.

[0033] Specifically, to solve the problem in the prior art that a single node cannot meet the graphics processor resource requirements of the inference model, resulting in the failure of the inference model deployment, the inference model is divided into multiple sub-inference models, and the graphics processor resource requirements of each sub-inference model are smaller. In some embodiments, as Figure 2 shown, the inference model includes 12 neural network layers, denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, CNN12 respectively. According to the segmentation information, CNN1, CNN2, and CNN3 are divided into one sub-inference model, CNN4, CNN5, and CNN6 are divided into one sub-inference model, CNN7, CNN8, and CNN9 are divided into one sub-inference model, and CNN10, CNN11, and CNN12 are divided into one sub-inference model.

[0034] Specifically, in some embodiments, the first Pod starts a program to load the inference model and loads the inference model into memory using keras of TensorFlow.

[0035] There are M pieces of the segmentation information, and step S101 can be implemented as:

[0036] The first Pod executes an extraction step, extracts the Nth sub-inference model from the inference model according to the Nth piece of segmentation information, and adds the Nth sub-inference model to the scheduling queue, where 1 ≤ N ≤ M, N is a positive integer natural number, and the Nth piece of segmentation information includes the name of the first neural network layer of the Nth sub-inference model and the name of the last neural network layer of the Nth sub-inference model.

[0037] The above-mentioned first Pod repeats the above-mentioned extraction step at least once until M of the above-mentioned sub-inference models are added to the above-mentioned scheduling queue, and the output of the previous sub-inference model in the above-mentioned scheduling queue is the input of the subsequent sub-inference model.

[0038] In this embodiment, in some implementation manners, there are 4 pieces of segmentation information, and the inference model includes 12 neural network layers, which are respectively denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, and CNN12. The first piece of segmentation information includes the names of CNN1 and CNN3, the second piece of segmentation information includes the names of CNN4 and CNN6, the third piece of segmentation information includes the names of CNN7 and CNN9, and the fourth piece of segmentation information includes the names of CNN10 and CNN12. First, the first Pod extracts CNN1, CNN2, and CNN3 from the inference model according to the first piece of segmentation information to obtain the first sub-inference model, and adds the first sub-inference model to the scheduling queue. After that, the first Pod extracts CNN4, CNN5, and CNN6 from the inference model according to the second piece of segmentation information to obtain the second sub-inference model, and adds the second sub-inference model to the scheduling queue. After that, the first Pod extracts CNN7, CNN8, and CNN9 from the inference model according to the third piece of segmentation information to obtain the third sub-inference model, and adds the third sub-inference model to the scheduling queue. Finally, the first Pod extracts CNN10, CNN11, and CNN12 from the inference model according to the fourth piece of segmentation information to obtain the fourth sub-inference model, and adds the fourth sub-inference model to the scheduling queue. Adding all 4 sub-inference models to the scheduling queue facilitates calling.

[0039] Step S102, the above-mentioned first Pod schedules the above-mentioned second Pod to the corresponding node and deploys the above-mentioned sub-inference model on the corresponding second Pod.

[0040] Among them, at least one graphics processor is configured on one of the above-mentioned nodes. The above-mentioned second Pods correspond to the above-mentioned nodes one by one, and the above-mentioned second Pods correspond to the above-mentioned sub-inference models one by one. The above-mentioned second Pod is used to run the corresponding above-mentioned sub-inference model and send the output data of the corresponding above-mentioned sub-inference model to the above-mentioned second Pod that runs the next above-mentioned sub-inference model, and the above-mentioned multiple sub-inference models run in a preset order.

[0041] Specifically, the graphics processor resource requirements of each sub-inference model are small, and each sub-inference model can find a node that meets the graphics processor resource requirements of the sub-inference model, thereby avoiding the situation where the inference model deployment fails when the graphics processor resources included in a node cannot meet the inference model, solving the problem in the prior art that a single node cannot meet the graphics processor resource requirements of the inference model and causing the inference model deployment to fail, and improving the utilization rate of the graphics processor resources distributed on multiple nodes in the cluster.

[0042] The above-mentioned inference model includes M above-mentioned sub-inference models, and the above-mentioned step S102 can be implemented as:

[0043] The above-mentioned first Pod executes a creation step to create the Nth above-mentioned second Pod, where 1 ≤ N ≤ M and N is a positive integer;

[0044] The above-mentioned first Pod executes a processing step to determine the Nth above-mentioned node according to the first quantity, and schedule the Nth above-mentioned second Pod to the Nth above-mentioned node. The above-mentioned first quantity is the number of the above-mentioned graphics processors that need to be called to run the Nth above-mentioned sub-inference model, and the above-mentioned first quantity is less than or equal to the second quantity, and the above-mentioned second quantity is the number of the above-mentioned graphics processors of the Nth above-mentioned node;

[0045] Specifically, the above-mentioned first quantity is obtained by the first Pod from the front end.

[0046] The above-mentioned first Pod executes a first acquisition step to obtain the Nth above-mentioned sub-inference model from the scheduling queue, and determine the model weight of the Nth above-mentioned sub-inference model and the network structure of the Nth above-mentioned sub-inference model according to the Nth above-mentioned sub-inference model. The output of the previous above-mentioned sub-inference model in the scheduling queue is the input of the subsequent above-mentioned sub-inference model;

[0047] The above-mentioned first Pod executes a first sending step to send the model weight of the Nth above-mentioned sub-inference model, the network structure of the Nth above-mentioned sub-inference model, and the IP address of the (N + 1)th above-mentioned sub-inference model to the Nth above-mentioned second Pod;

[0048] The above-mentioned first Pod executes a start step to start the Nth above-mentioned inference model;

[0049] The above-mentioned first Pod repeatedly executes the above-mentioned creation step, the above-mentioned processing step, the above-mentioned first acquisition step, the above-mentioned first sending step, and the above-mentioned start step at least once until the start of M above-mentioned inference models is completed.

[0050] In this embodiment, as Figure 3As shown, in some embodiments, there are 4 pieces of segmentation information, and the inference model includes 12 neural network layers, denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, and CNN12 respectively. There are 4 sub-inference models in the scheduling queue, from the head to the tail of the queue are the 1st sub-inference model (including CNN1, CNN2, CNN3), the 2nd sub-inference model (including CNN4, CNN5, CNN6), the 3rd sub-inference model (including CNN7, CNN8, CNN9), and the 4th sub-inference model (including CNN10, CNN11, CNN12). First, the first Pod deploys the 4th sub-inference model. The first Pod creates the 4th second Pod, records the IP address of the 4th second Pod, determines the 4th node according to the number of graphics processors required by the 4th sub-inference model (the number of graphics processors of the 4th node is greater than or equal to the number of graphics processors required by the 4th sub-inference model), schedules the 4th second Pod to the 4th node. After the 4th second Pod is scheduled, the 4th second Pod starts the parameter receiving program. The first Pod reads the 4th sub-inference model from the end of the queue, and sends the model weights, network structure, and the IP address of the first Pod to the 4th second Pod. After the parameter receiving program of the 4th second Pod receives the model weights and network structure of the 4th sub-inference model, it starts the Runtime of the 4th sub-inference model to complete the deployment of the 4th sub-inference model. Then, the first Pod deploys the 3rd sub-inference model. The process is as follows: The first Pod creates the 3rd second Pod, records the IP address of the 3rd second Pod, determines the 3rd node according to the number of graphics processors required by the 3rd sub-inference model (the number of graphics processors of the 3rd node is greater than or equal to the number of graphics processors required by the 3rd sub-inference model), schedules the 3rd second Pod to the 3rd node. After the 3rd second Pod is scheduled, the 3rd second Pod starts the parameter receiving program. The first Pod reads the 3rd sub-inference model from the scheduling queue, and sends the model weights, network structure, and the IP address of the 4th second Pod to the 3rd second Pod. After the parameter receiving program of the 3rd second Pod receives the model weights and network structure of the 3rd sub-inference model, it starts the Runtime of the 3rd sub-inference model to complete the deployment of the 3rd sub-inference model. After that, the first Pod deploys the 2nd sub-inference model. The process is as follows: The first Pod creates the 2nd second Pod, records the IP address of the 2nd second Pod, determines the 2nd node according to the number of graphics processors required by the 2nd sub-inference model (the number of graphics processors of the 2nd node is greater than or equal to the number of graphics processors required by the 2nd sub-inference model),Schedule the second second Pod to the second node. After the second second Pod is scheduled, the second second Pod starts the parameter receiving program. The first Pod reads the second sub-inference model from the scheduling queue, and sends the model weights, network structure of the second sub-inference model, and the IP address of the third second Pod to the second second Pod. After the parameter receiving program of the second second Pod receives the model weights and network structure of the second sub-inference model, it starts the Runtime of the second sub-inference model to complete the deployment of the second sub-inference model. Finally, the first Pod deploys the first sub-inference model. The process is as follows: The first Pod creates the first second Pod, records the IP address of the first second Pod, determines the first node according to the number of GPUs required by the first sub-inference model (the number of GPUs of the first node is greater than or equal to the number of GPUs required by the first sub-inference model), and schedules the first second Pod to the first node. After the first second Pod is scheduled, the first second Pod starts the parameter receiving program. The first Pod reads the first sub-inference model from the scheduling queue, and sends the model weights, network structure of the first sub-inference model, and the IP address of the second second Pod to the first second Pod. After the parameter receiving program of the first second Pod receives the model weights and network structure of the first sub-inference model, it starts the Runtime of the first sub-inference model to complete the deployment of the first sub-inference model.,

[0051] Step S103: The first Pod sends the input data of the inference model to the second Pod running the first sub-inference model, and obtains the output data of the inference model from the second Pod running the last sub-inference model.

[0052] Specifically, each sub-inference model runs a part of the inference, and all sub-inference models together complete the inference process of the entire inference model. In some embodiments, the inference model includes 12 neural network layers, denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, and CNN12 respectively. There are 4 sub-inference models, namely the first sub-inference model (including CNN1, CNN2, and CNN3), the second sub-inference model (including CNN1, CNN2, and CNN3), the third sub-inference model (including CNN7, CNN8, and CNN9), and the fourth sub-inference model (including CNN10, CNN11, and CNN12). The first sub-inference model, the second sub-inference model, the third sub-inference model, and the fourth sub-inference model are run in sequence.

[0053] In an alternative solution, after the above step S103, the method further includes:

[0054] The first Pod sends multiple first data packets to the second Pod running the first sub-inference model, and obtains multiple second data packets from the second Pod running the last sub-inference model. The first data packet includes the input data and identifier of the inference model. The second data packet includes the output data and identifier of the inference model. One first data packet corresponds to one input data of the inference model. One second data packet corresponds to one output data of the inference model. One input data of the inference model corresponds to one identifier, and one input data of the inference model corresponds to one output data of the inference model.

[0055] In this embodiment, in some implementations, the inference model needs to perform inferences on multiple input data. The inference model performs an inference on one input data to obtain one output data. To avoid confusion about the output data corresponding to the input data, an identifier is bound to each input data. After the inference process ends and the output data is obtained, the input data corresponding to each output data is determined according to the identifier.

[0056] The backend is configured with a database for storing multiple input data of the inference model. The first Pod is communicatively connected to the database. In an alternative solution, after step S103, the method further includes:

[0057] The first Pod obtains a target identifier and target output data. The target identifier is the identifier in the target data packet, and the target output data is the output data of the inference model in the target data packet. The target data packet is one of the multiple second data packets.

[0058] The first Pod obtains target input data from the database according to the target identifier and the first mapping relationship. The target input data is the input data corresponding to the target identifier. The first mapping relationship is the mapping relationship between the identifier and the input data of the inference model.

[0059] The first Pod sends the target input data and the target output data to the front end.

[0060] In this embodiment, in some implementation manners, the inference model needs to perform inferences on multiple input data. The inference model performs an inference on one input data to obtain one output data. To avoid confusing the output data corresponding to the input data, an identifier is bound to each input data. After the output data and the identifier are obtained at the end of the inference process, the input data corresponding to each output data is determined according to the identifier, and each output data and its corresponding input data are sent to the front end.

[0061] Through the above steps, the inference model is divided into multiple sub-inference models. Each sub-inference model has a relatively small demand for graphics processor resources. Each sub-inference model can find a node that meets the graphics processor resource requirements of the sub-inference model. Each sub-inference model runs a part of the inference, and all sub-inference models jointly complete the inference process of the entire inference model. Furthermore, it avoids the situation where the inference model deployment fails when the graphics processor resources included in a node cannot meet the requirements of the inference model, solves the problem in the prior art that a single node cannot meet the graphics processor resource requirements of the inference model resulting in the failure of inference model deployment, and improves the utilization rate of the graphics processor resources distributed on multiple nodes in the cluster.

[0062] The inference model is divided into M sub-inference models. A first Pod and multiple second Pods are configured in the backend. The Nth second Pod among the multiple second Pods is the second Pod that runs the Nth sub-inference model. The sub-inference models run in a preset order, and N is a positive integer natural number.

[0063] In this embodiment, a method running on the multiple second Pods is provided. Figure 4 It is a flowchart according to an embodiment of the present application, as Figure 4 shown. The process includes the following steps:

[0064] Step S201, the first second Pod among the multiple second Pods receives the input data of the inference model from the first Pod;

[0065] Specifically, as Figure 5 shown, the front end sends the input data of the inference model to the first Pod, and the first Pod sends the input data of the inference model to the first second Pod.

[0066] Step S202, the first second Pod among the multiple second Pods inputs the input data of the inference model into the first sub-inference model to obtain the output data of the first sub-inference model;

[0067] Step S203: The first second Pod among the multiple second Pods sends the output data of the first sub-inference model to the second second Pod according to the IP address of the second second Pod.

[0068] Step S204: First receiving step. The K-th second Pod among the multiple second Pods receives the output data of the (K - 1)-th sub-inference model from the (K - 1)-th second Pod, where 2 ≤ K ≤ M and K is a positive integer.

[0069] Step S205: First control step. The K-th second Pod among the multiple second Pods inputs the output data of the (K - 1)-th sub-inference model into the K-th sub-inference model to obtain the output data of the K-th sub-inference model.

[0070] Step S206: Second sending step. The K-th second Pod among the multiple second Pods sends the output data of the K-th sub-inference model to the (K + 1)-th second Pod according to the IP address of the (K + 1)-th second Pod.

[0071] Step S207: First determination step. The K-th second Pod among the multiple second Pods determines whether the IP address of the (K + 1)-th second Pod is the IP address of the first Pod.

[0072] Step S208: Repeat the above first receiving step, first control step, second sending step, and first determination step at least once until the IP address of the (K + 1)-th second Pod is the IP address of the first Pod and then end.

[0073] In this embodiment, in some implementation manners, the inference model includes 12 neural network layers, denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, and CNN12 respectively. There are 4 sub-inference models, namely the first sub-inference model (including CNN1, CNN2, and CNN3), the second sub-inference model (including CNN4, CNN5, and CNN6), the third sub-inference model (including CNN7, CNN8, and CNN9), and the fourth sub-inference model (including CNN10, CNN11, and CNN12), as Figure 5As shown, after the first second Pod receives the input data of the inference model from the first Pod, the first second Pod inputs the input data of the inference model into the first sub-inference model (including CNN1, CNN2, and CNN3) to obtain the output data of the first sub-inference model. The first second Pod sends the output data of the first sub-inference model to the second second Pod according to the IP address of the second second Pod. The second second Pod receives the output data of the first sub-inference model from the first second Pod. The second second Pod inputs the output data of the first sub-inference model into the second sub-inference model (including CNN4, CNN5, and CNN6) to obtain the output data of the second sub-inference model. The second second Pod sends the output data of the second sub-inference model to the third second Pod according to the IP address of the third second Pod. The third second Pod receives the output data of the second sub-inference model from the second second Pod. The third second Pod inputs the output data of the second sub-inference model into the third sub-inference model (including CNN7, CNN8, and CNN9) to obtain the output data of the third sub-inference model. The third second Pod sends the output data of the third sub-inference model to the fourth second Pod according to the IP address of the fourth second Pod. The fourth second Pod receives the output data of the third sub-inference model from the third second Pod. The fourth second Pod inputs the output data of the third sub-inference model into the fourth sub-inference model (including CNN10, CNN11, and CNN12) to obtain the output data of the fourth sub-inference model. The fourth second Pod sends the output data of the fourth sub-inference model to the first Pod according to the IP address of the first Pod, and the first Pod sends the output data of the inference model to the front end.

[0074] In an alternative solution, the above method further includes:

[0075] Step S301, the first second Pod among the multiple second Pods receives a third data packet from the first Pod, and the third data packet includes: the input data of the inference model and an identifier;

[0076] Specifically, the identifier is generated by the first Pod.

[0077] Step S302, the first second Pod among the multiple second Pods inputs the input data of the inference model into the first sub-inference model to obtain the output data of the first sub-inference model;

[0078] Step S303, the first second Pod among the multiple second Pods packs the output data of the first sub-inference model and the identifier into a fourth data packet;

[0079] Step S304. The first second Pod among the multiple second Pods sends the fourth data packet to the second second Pod according to the address of the second second Pod.

[0080] Step S305. Second receiving step. The K-th second Pod among the multiple second Pods receives the fifth data packet from the (K - 1)-th second Pod. The fifth data packet includes the output data of the (K - 1)-th sub-inference model and the identifier.

[0081] Step S306. Second control step. The K-th second Pod among the multiple second Pods inputs the output data of the (K - 1)-th sub-inference model into the K-th sub-inference model to obtain the output data of the K-th sub-inference model.

[0082] Step S307. Packing step. The K-th second Pod among the multiple second Pods packs the output data of the K-th sub-inference model and the identifier into a sixth data packet.

[0083] The above Step S307 can be implemented as:

[0084] The K-th second Pod among the multiple second Pods compresses the output data of the K-th sub-inference model, and packs the compressed output data of the K-th sub-inference model and the identifier into the above sixth data packet.

[0085] In this embodiment, to improve the running speed of the inference model, the K-th second Pod compresses the output data of the K-th sub-inference model first, and then packs the output data of the K-th sub-inference model and the identifier into a sixth data packet, thereby accelerating the data transmission rate and improving the running speed of the inference model.

[0086] Step S308. Third sending step. The K-th second Pod among the multiple second Pods sends the sixth data packet to the (K + 1)-th second Pod according to the address of the (K + 1)-th second Pod.

[0087] Step S309. Second determination step. The K-th second Pod among the multiple second Pods determines whether the IP address of the (K + 1)-th second Pod is the IP address of the first Pod.

[0088] Step S310. Repeat the above second receiving step, the above second control step, the above packing step, the above third sending step, and the above second determination step at least once until the IP address of the (K + 1)-th second Pod is the IP address of the first Pod and then end.

[0089] In this embodiment, in some embodiments, the inference model needs to perform inferences on multiple input data. The inference model performs an inference on one input data to obtain one output data. To avoid confusion about the output data corresponding to the input data, an identifier is bound to each input data. After the inference process ends and the output data is obtained, the input data corresponding to each output data is determined according to the identifier.

[0090] In this embodiment, in some embodiments, the inference model includes 12 neural network layers, denoted as CNN1, CNN2, CNN3, CNN4, CNN5, CNN6, CNN7, CNN8, CNN9, CNN10, CNN11, and CNN12 respectively. There are 4 sub-inference models, namely the first sub-inference model (including CNN1, CNN2, CNN3), the second sub-inference model (including CNN4, CNN5, CNN6), the third sub-inference model (including CNN7, CNN8, CNN9), and the fourth sub-inference model (including CNN10, CNN11, CNN12), as Figure 5As shown, after the first second Pod receives the input data and identifier of the inference model from the first Pod, the first second Pod inputs the input data of the inference model into the first sub-inference model (including CNN1, CNN2, and CNN3) to obtain the output data of the first sub-inference model. The first second Pod packs the output data and identifier of the first sub-inference model into a data packet and sends the data packet to the second second Pod according to the IP address of the second second Pod. The second second Pod receives the data packet. The second second Pod inputs the output data of the first sub-inference model into the second sub-inference model (including CNN4, CNN5, and CNN6) to obtain the output data of the second sub-inference model. The second second Pod packs the output data and identifier of the second sub-inference model into a data packet and sends the data packet to the third second Pod according to the IP address of the third second Pod. The third second Pod receives the data packet. The third second Pod inputs the output data of the second sub-inference model into the third sub-inference model (including CNN7, CNN8, and CNN9) to obtain the output data of the third sub-inference model. The third second Pod packs the output data and identifier of the third sub-inference model into a data packet and sends the data packet to the fourth second Pod according to the IP address of the fourth second Pod. The fourth second Pod receives the data packet. The fourth second Pod inputs the output data of the third sub-inference model into the fourth sub-inference model (including CNN10, CNN11, and CNN12) to obtain the output data of the fourth sub-inference model. The fourth second Pod packs the output data and identifier of the fourth sub-inference model into a data packet and sends the data packet to the first Pod according to the IP address of the first Pod. The first Pod receives the data packet, and the first Pod sends the output data of the inference model to the front end.

[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0092] In this embodiment, a backend is further provided. The backend is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. Although the backend described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated. The backend includes:

[0093] A first Pod, and the first Pod is communicatively connected to the front end;

[0094] A second Pod;

[0095] Wherein,

[0096] When the first Pod receives the segmentation information from the front end and the inference model, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information. The segmentation information corresponds to the sub-inference models one by one. One sub-inference model includes multiple neural network layers. One piece of segmentation information includes the name of the first neural network layer of the corresponding sub-inference model and the name of the last neural network layer of the corresponding sub-inference model;

[0097] The first Pod schedules the second Pod to the corresponding node and deploys the sub-inference model on the corresponding second Pod. At least one graphics processor is configured on one node. The second Pod corresponds to the node one by one, and the second Pod corresponds to the sub-inference model one by one. The second Pod is used to run the corresponding sub-inference model and send the output data of the corresponding sub-inference model to the second Pod that runs the next sub-inference model. The multiple sub-inference models run in a preset order;

[0098] The first Pod sends the input data of the inference model to the second Pod that runs the first sub-inference model and obtains the output data of the inference model from the second Pod that runs the last sub-inference model.

[0099] Through the above steps, the inference model is divided into multiple sub-inference models. The graphics processor resource requirements of each sub-inference model are small. Each sub-inference model can find a node that meets the graphics processor resource requirements of the sub-inference model. Each sub-inference model runs a part of the inference, and all sub-inference models jointly complete the inference process of the entire inference model. Furthermore, it avoids the situation where the inference model deployment fails when the graphics processor resources included in one node cannot meet the inference model, solves the problem in the prior art that the graphics processor resources of a single node cannot meet the requirements of the inference model and leads to the failure of the inference model deployment, and improves the utilization rate of the graphics processor resources distributed on multiple nodes in the cluster.

[0100] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.

[0101] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0102] In an exemplary embodiment, the above computer-readable storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0103] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0104] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0105] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0106] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.

[0107] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for running an inference model, characterized in that, a first Pod and multiple second Pods are configured in the backend, the first Pod is communicatively connected to the frontend, the method is applied to the first Pod, the inference model includes multiple neural network layers, and the neural network layers are connected in sequence. The method includes: When the first Pod receives the segmentation information and the inference model from the frontend, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information. The segmentation information corresponds to the sub-inference models one by one. One sub-inference model includes multiple neural network layers. One piece of segmentation information includes the name of the first neural network layer and the name of the last neural network layer of the corresponding sub-inference model; The first Pod schedules the second Pods to the corresponding nodes and deploys the sub-inference models in the corresponding second Pods. At least one graphics processor is configured on one node. The second Pods correspond to the nodes one by one, and the second Pods correspond to the sub-inference models one by one. The second Pod is used to run the corresponding sub-inference model and send the output data of the corresponding sub-inference model to the second Pod that runs the next sub-inference model, and the multiple sub-inference models run in a preset order; The first Pod sends the input data of the inference model to the second Pod that runs the first sub-inference model and obtains the output data of the inference model from the second Pod that runs the last sub-inference model.

2. The method according to claim 1, characterized in that, there are M pieces of segmentation information. When the first Pod receives the segmentation information and the inference model from the frontend, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information, including: The first Pod executes an extraction step. According to the Nth piece of segmentation information, the Nth sub-inference model is extracted from the inference model and the Nth sub-inference model is added to the scheduling queue, where 1 ≤ N ≤ M and N is a positive integer natural number. The Nth piece of segmentation information includes the name of the first neural network layer and the name of the last neural network layer of the Nth sub-inference model; The first Pod repeats the extraction step at least once until M sub-inference models are added to the scheduling queue. The output of the previous sub-inference model in the scheduling queue is the input of the next sub-inference model.

3. The method according to claim 1, characterized in that, the inference model includes M sub-inference models. The first Pod schedules the second Pods to the corresponding nodes and deploys the sub-inference models in the corresponding second Pods, including: The first Pod executes a creation step to create the Nth second Pod, where 1 ≤ N ≤ M, N is a positive integer natural number; The first Pod executes a processing step, determines the Nth node according to a first quantity, and schedules the Nth second Pod to the Nth node. The first quantity is the number of graphics processors required to run the Nth sub-inference model, and the first quantity is less than or equal to a second quantity, where the second quantity is the number of graphics processors of the Nth node; The first Pod executes a first obtaining step, obtains the Nth sub-inference model from a scheduling queue, and determines the model weights of the Nth sub-inference model and the network structure of the Nth sub-inference model according to the Nth sub-inference model. The output of the previous sub-inference model in the scheduling queue is the input of the subsequent sub-inference model; The first Pod executes a first sending step, and sends the model weights of the Nth sub-inference model, the network structure of the Nth sub-inference model, and the IP address of the (N + 1)th sub-inference model to the Nth second Pod; The first Pod executes a starting step, and starts the Nth inference model; The first Pod repeatedly executes the creating step, the processing step, the first obtaining step, the first sending step, and the starting step at least once until the starting of M inference models is completed.

4. The method according to claim 1, wherein, after the first Pod schedules the second Pod to the corresponding node and deploys the sub-inference model on the corresponding second Pod, the method further includes: The first Pod sends a plurality of first data packets to the second Pod running the first sub-inference model, and obtains a plurality of second data packets from the second Pod running the last sub-inference model. The first data packets include the input data and identifiers of the inference model, and the second data packets include the output data and identifiers of the inference model. One first data packet corresponds to one input data of the inference model, one second data packet corresponds to one output data of the inference model, one input data of the inference model corresponds to one identifier, and one input data of the inference model corresponds to one output data of the inference model.

5. A method for running an inference model, wherein, the inference model is divided into M sub-inference models, a first Pod and a plurality of second Pods are configured in the backend, and the Nth second Pod among the plurality of second Pods is the second Pod running the Nth sub-inference model. The sub-inference models run in a preset order, and N is a positive integer natural number. The method is applied to the plurality of second Pods, and the method includes: The first second Pod among the plurality of second Pods receives the input data of the inference model from the first Pod; The first second Pod among the plurality of second Pods inputs the input data of the inference model into the first sub-inference model to obtain the output data of the first sub-inference model; The first second Pod among the multiple second Pods sends the output data of the first sub-inference model to the second second Pod according to the IP address of the second second Pod; First receiving step, the K-th second Pod among the multiple second Pods receives the output data of the (K - 1)-th sub-inference model from the (K - 1)-th second Pod, where 2 ≤ K ≤ M and K is a positive integer; First control step, the K-th second Pod among the multiple second Pods inputs the output data of the (K - 1)-th sub-inference model into the K-th sub-inference model to obtain the output data of the K-th sub-inference model; Second sending step, the K-th second Pod among the multiple second Pods sends the output data of the K-th sub-inference model to the (K + 1)-th second Pod according to the IP address of the (K + 1)-th second Pod; First determining step, the K-th second Pod among the multiple second Pods determines whether the IP address of the (K + 1)-th second Pod is the IP address of the first Pod; Repeat the first receiving step, the first control step, the second sending step, and the first determining step at least once until the IP address of the (K + 1)-th second Pod is the IP address of the first Pod and then end.

6. The method according to claim 5, wherein, the method further includes: the first second Pod among the multiple second Pods receives a third data packet from the first Pod, and the third data packet includes: the input data of the inference model and an identifier; the first second Pod among the multiple second Pods inputs the input data of the inference model into the first sub-inference model to obtain the output data of the first sub-inference model; the first second Pod among the multiple second Pods packs the output data of the first sub-inference model and the identifier into a fourth data packet; the first second Pod among the multiple second Pods sends the fourth data packet to the second second Pod according to the address of the second second Pod; Second receiving step, the K-th second Pod among the multiple second Pods receives a fifth data packet from the (K - 1)-th second Pod, and the fifth data packet includes the output data of the (K - 1)-th sub-inference model and the identifier; Second control step, the K-th second Pod among the multiple second Pods inputs the output data of the (K - 1)-th sub-inference model into the K-th sub-inference model to obtain the output data of the K-th sub-inference model; Packing step, the K-th second Pod among the multiple second Pods packs the output data of the K-th sub-inference model and the identifier into a sixth data packet; Third sending step, the K-th second Pod among the multiple second Pods sends the sixth data packet to the (K + 1)-th second Pod according to the address of the (K + 1)-th second Pod; Second determination step, the K-th second Pod among the multiple second Pods determines whether the IP address of the (K + 1)-th second Pod is the IP address of the first Pod; Repeat the second receiving step, the second control step, the packing step, the third sending step, and the second determination step at least once until the IP address of the (K + 1)-th second Pod is the IP address of the first Pod and then end.

7. The method according to claim 6, wherein, The K-th second Pod among the multiple second Pods packs the output data of the K-th sub-inference model and the identifier into a sixth data packet, including: The K-th second Pod among the multiple second Pods compresses the output data of the K-th sub-inference model, and packs the compressed output data of the K-th sub-inference model and the identifier into the sixth data packet.

8. A backend, wherein, The inference model includes multiple neural network layers which are connected in sequence, and the backend includes: A first Pod which is communicatively connected to the frontend; A second Pod; wherein, When the first Pod receives the segmentation information and the inference model from the frontend, the first Pod divides the inference model into multiple sub-inference models according to each piece of segmentation information, the segmentation information corresponds to the sub-inference model one by one, one sub-inference model includes multiple neural network layers, and one piece of segmentation information includes the name of the first neural network layer of the corresponding sub-inference model and the name of the last neural network layer of the corresponding sub-inference model; The first Pod schedules the second Pod to the corresponding node, and deploys the sub-inference model on the corresponding second Pod. One node is configured with at least one graphics processor, the second Pod corresponds to the node one by one, the second Pod corresponds to the sub-inference model one by one, the second Pod is used to run the corresponding sub-inference model and send the output data of the corresponding sub-inference model to the second Pod running the next sub-inference model, and the multiple sub-inference models run in a preset order; The first Pod sends the input data of the inference model to the second Pod running the first sub-inference model, and obtains the output data of the inference model from the second Pod running the last sub-inference model.

9. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 4, or implements the steps of the method described in any one of claims 5 to 7.

10. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 4, or implements the steps of the method described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Task processing method and device and storage medium

    CN112148348A

  • Optimization method and device of inference model, electronic equipment and storage medium

    CN116227599A