Inference method, system and storage medium of a computation graph
By partitioning and recursively transforming the original type computation graph, generating a file to be inferred, and sending it to a remote device, the problem of dynamic computation graphs being unable to be converted into static computation graphs is solved, thus improving the inference efficiency of computation graphs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2023-03-16
- Publication Date
- 2026-07-31
AI Technical Summary
During remote inference, dynamic computation graphs cannot be successfully converted into static computation graphs, resulting in low efficiency in remote inference of computation graphs. Existing technologies have not been able to effectively solve this problem.
By monitoring the primitive type computation graph to be inferred, when the response conversion fails, it is divided into a first-level sub-primitive type computation graph, and then converted into a target type computation graph through recursion. The resulting inference file is sent to a remote device for inference, thus avoiding human modification.
This achieves the goal of converting any primitive type computation graph into a target type computation graph, improving the inference efficiency of computation graphs and solving the problem of low inference efficiency in computation graphs.
Smart Images

Figure CN116415666B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing, and more specifically, to a reasoning method, system, and storage medium for a computation graph. Background Technology
[0002] Currently, remote inference typically requires converting a dynamic computation graph into a static computation graph, then exporting the corresponding model file, and finally performing remote inference on that model file. However, in many scenarios, the dynamic computation graph fails to convert successfully to a static computation graph, so users usually need to manually modify the static computation graph to ensure a successful conversion.
[0003] Therefore, in many scenarios, dynamic computation graphs cannot be directly converted into static computation graphs when performing remote inference, which easily leads to the technical problem of low efficiency in remote inference of computation graphs.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a reasoning method, system, and storage medium for computation graphs, to at least solve the technical problem of low reasoning efficiency in computation graphs.
[0006] According to one aspect of the embodiments of this application, a reasoning method for a computation graph is provided. The method may include: monitoring a primitive type computation graph to be reasoned, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to a detected failure to convert a primitive type computation graph into a corresponding target type computation graph, dividing the primitive type computation graph into first-level sub-primary type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; converting the first-level sub-primary type computation graphs into at least one corresponding target type computation graph; generating a reasoning file from the target type computation graph corresponding to the first-level sub-primary type computation graph, and sending the reasoning file to a remote device, wherein the reasoning file is a file provided to the remote device for reasoning.
[0007] According to another aspect of the embodiments of this application, an inference apparatus for a computation graph is also provided. The apparatus may include: a monitoring module for monitoring a primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; a partitioning module for partitioning the primitive type computation graph into first-level sub-primitive type computation graphs in response to a detected failure to convert the primitive type computation graph into a corresponding target type computation graph, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; a conversion module for converting the first-level sub-primitive type computation graphs into at least one corresponding target type computation graph; and a sending module for generating a file to be inferred from the target type computation graph corresponding to the first-level sub-primitive type computation graph and sending the file to be inferred to a remote device, wherein the file to be inferred is a file provided to the remote device for inference.
[0008] According to another aspect of the embodiments of this application, a computation graph inference system is also provided, comprising: a local device for monitoring a primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to the detected failure to convert the primitive type computation graph into a corresponding target type computation graph, dividing the primitive type computation graph into first-level sub-primary type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; converting the first-level sub-primary type computation graphs into at least one corresponding target type computation graph; generating a file to be inferred from the target type computation graph corresponding to the first-level sub-primary type computation graph; and a remote device for acquiring the file to be inferred and performing inference on the file to be inferred to obtain an inference result.
[0009] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is run by a processor, it controls the device where the computer-readable storage medium is located to execute a reasoning method of a computation graph.
[0010] In this embodiment, the original type computation graph is transformed. When the transformation from the original type computation graph to the corresponding target type computation graph fails, the original type computation graph can be divided into first-level sub-original type computation graphs. These first-level sub-original type computation graphs are then transformed into at least one corresponding target type computation graph. Finally, a file to be inferred is generated from the target type computation graph corresponding to the first-level sub-original type computation graph, and this file is sent to a remote device so that the remote device can perform inference operations based on the file. In other words, in this embodiment, when the transformation from the original type computation graph to the corresponding target type computation graph fails, the original type computation graph can be divided into sub-original type computation graphs, and these sub-original type computation graphs can be further transformed. This process is repeated recursively until the transformation to the target type computation graph is successfully completed. This allows any original type computation graph to be transformed into a target type computation graph without manual modification of the original type computation graph. The operation is relatively simple, achieving the goal of transforming any original type computation graph into a target type computation graph for inference. This improves the inference efficiency of computation graphs and solves the technical problem of low inference efficiency.
[0011] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0013] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a computation graph reasoning method according to an embodiment of this application;
[0014] Figure 2 This is a structural block diagram of a computational environment for a computational graph reasoning method according to an embodiment of this application;
[0015] Figure 3 This is a structural block diagram of a service mesh according to an embodiment of this application;
[0016] Figure 4 This is a flowchart of a computation graph reasoning method according to an embodiment of this application;
[0017] Figure 5 This is a schematic diagram of a computational graph reasoning method according to an embodiment of this application;
[0018] Figure 6 This is a flowchart of another computational graph reasoning method according to an embodiment of this application;
[0019] Figure 7 This is a schematic diagram of a computation graph inference system according to an embodiment of this application;
[0020] Figure 8 This is a structural block diagram of a computation graph inference device according to an embodiment of this application;
[0021] Figure 9 This is a structural block diagram of a computer terminal with a computational graph according to an embodiment of this application. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0025] A computation graph is a directed acyclic graph used to describe computations.
[0026] Primitive type computation graphs, where the construction and computation of the computation graph occur simultaneously, such as dynamic computation graphs;
[0027] The target type computation graph separates the construction and computation of the computation graph, and predefines the entire operation flow, for example, a static computation graph;
[0028] First-level primitive type computation graph, and subgraphs of the primitive type computation graph;
[0029] Second-level subprimitive type computation graph, and subgraphs of the first-level subprimitive type computation graph;
[0030] The target operator is used to build the basic unit of the primitive type computation graph;
[0031] Local inference involves calling the application programming interface (API) provided by the deep learning framework on a single machine to perform computations on the constructed computation graph, and the inference is also executed on that machine.
[0032] Remote inference involves users calling APIs provided by a remote inference framework on their local machines to perform inference functions, while the remote machine actually executes the inference. The execution process of remote inference is as follows: the local machine sends the computation graph to the remote machine via the network based on the remote inference framework, and also sends the user's inference request to the remote machine via the network. The remote machine completes the inference operation based on the computation graph and the user's inference request, and returns the inference result to the local machine via the network.
[0033] Example 1
[0034] According to an embodiment of this application, a reasoning method for a computation graph is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0035] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a computation graph reasoning method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor (MCU) or a field-programmable gate array (FPGA), etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output (I / O) interface, a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0036] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0037] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the computation graph inference method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned computation graph inference method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0039] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0040] Figure 1The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned computer terminal 10 (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a computing node in computing environment 201. Figure 2 A block diagram of the computational environment for an alternative computational graph reasoning method is shown, such as... Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (represented as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0041] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0042] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0043] In one embodiment based on container virtualization, several containers of a service can be assembled into a workload container group (Pod) (e.g., a Kubernetes Pod). For example, such as Figure 2As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.
[0044] During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201, and executing one or more functions of one service may require invoking one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.
[0045] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0046] In another alternative embodiment, Figure 3 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a service mesh. Figure 3 A structural block diagram of a service mesh is shown, such as Figure 3 As shown, the service mesh 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines.
[0047] like Figure 3 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 300. In one implementation, application service instance A runs as a container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on machine / workload container group 316 (Pod).
[0048] In one implementation, application service instance A can be a product query service implemented by performing remote inference on a model file corresponding to a static computation graph, and application service instance B can be a product order placement service implemented by performing remote inference on a model file corresponding to another static computation graph. The model file corresponding to the aforementioned static computation graph can be provided by the embodiments of this application. Figure 4 The reasoning method for the computational graph shown is obtained.
[0049] like Figure 3 As shown, application service instance A and grid agent (sidecar) 303 coexist in machine workload container group 614, and application service instance B and grid agent 305 coexist in machine workload container 314. Grid agents 303 and 305 form the data plane layer of service mesh 300. Grid agents 303 and 305 run as containers / processes 304 and 306 respectively, and can receive requests 312 for product query services. Grid agent 303 and application service instance A can communicate bidirectionally, and grid agent 305 and application service instance B can also communicate bidirectionally. Furthermore, grid agents 303 and 305 can also communicate bidirectionally with each other.
[0050] In one implementation, traffic from application service instance A is routed to the appropriate destination via mesh proxy 303, and network traffic from application service instance B is routed to the appropriate destination via mesh proxy 305. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (Google Remote Procedure Call, gRPC), and open-source in-memory data structure storage systems (such as Remote Dictionary Server, Redis).
[0051] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the proxy (Envoy) in service mesh 300. The service mesh proxy configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0052] like Figure 3 As shown, the service mesh 300 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 301 within machine / workload container groups (machine / Pods) 302. Figure 3 As shown, the managed control plane component 301 communicates bidirectionally with grid agents 303 and 305. The managed control plane component 301 is configured to perform various control and management functions. For example, it receives telemetry data transmitted from grid agents 303 and 305 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 301 can also provide a user-facing application programming interface (API) to facilitate manipulation of network behavior and the provision of configuration data to grid agents 303 and 305.
[0053] Under the aforementioned operating environment, this application provides the following: Figure 4 The inference method for the computational graph shown is illustrated. It should be noted that the inference method for the computational graph in this embodiment can be derived from... Figure 1 The computer terminal execution of the illustrated embodiment can be applied to local devices. Figure 4 This is a flowchart of a computational graph reasoning method according to Embodiment 1 of this application. For example... Figure 4 As shown, the method may include the following steps:
[0054] Step S401: Monitor the original type computation graph to be reasoned, wherein the original type computation graph is a computation graph constructed and computed within the same time period.
[0055] In the technical solution provided by step S401 of this application, the original type computation graph to be reasoned can be monitored. The original type computation graph is a computation graph that is constructed and computed within the same time period. For example, the original type computation graph can be a dynamic computation graph, which can be called a dynamic graph.
[0056] Optionally, after detecting the computation graph of the original type to be inferred, the original type computation graph can be directly converted into a target type computation graph using the API provided by the deep learning framework. The target type computation graph can be a static computation graph.
[0057] In step S402, in response to the detected failure to convert the original type computation graph into the corresponding target type computation graph, the original type computation graph is divided into first-level sub-original type computation graphs, wherein the target type computation graph is a computation graph that is constructed and computed in different time periods.
[0058] In the technical solution provided by step S402 of this application, after monitoring the primitive type computation graph to be reasoned, a type conversion can be performed on the monitored primitive type computation graph, for example, a type conversion from a dynamic computation graph to a static computation graph. In response to the failure to convert the monitored primitive type computation graph to the corresponding target type computation graph, the primitive type computation graph can be divided into first-level sub-primitive type computation graphs.
[0059] In this embodiment, since there may be situations in many scenarios where the primitive type computation graph cannot be directly converted to the target type computation graph—for example, the primitive type computation graph contains syntax specific to Python, a computer programming language that allows for simple and efficient object-oriented programming, or the input and output of the primitive type computation graph are not types supported by the deep learning framework—the API provided by the deep learning framework will automatically throw an exception during the conversion process. When an exception is detected from the API provided by the deep learning framework, i.e., in response to a failure to convert the primitive type computation graph to the corresponding target type computation graph, the primitive type computation graph can then be divided into first-level sub-primitive type computation graphs for conversion.
[0060] Optionally, the aforementioned target type computation graph is a computation graph constructed and computed within different time periods, and can be a static computation graph, also referred to as a static graph. Since the original type computation graph can be a dynamic computation graph, this first-level sub-original type computation graph can be a first-level sub-dynamic computation graph obtained by partitioning the dynamic computation graph.
[0061] Optionally, if the primitive type computation graph conforms to the standard format of a deep learning framework, then the type and number of the first-level sub-primitive type computation graphs obtained after partitioning the primitive type computation graph are determined. For example, if the primitive type computation graph conforms to the standard format of a deep learning framework, then the first-level sub-primitive type computation graphs obtained after partitioning the primitive type computation graph can be as follows: Figure 5 Sub-dynamic calculation shown Figure 1 Sub-dynamic computation Figure 2 and sub-dynamic computation Figure 3 The examples provided are merely illustrative and do not constitute a limitation on the embodiments of this application.
[0062] Step S403: Transform the first-level subprime type computation graph into a corresponding at least one target type computation graph.
[0063] In the technical solution provided by step S403 of this application, the first-level sub-original type computation graph can be transformed into at least one corresponding target type computation graph. For example, the sub-dynamic computation graph can be transformed into at least one corresponding sub-static computation graph.
[0064] In this embodiment, after dividing the primitive type computation graph into first-level sub-primitive type computation graphs, the first-level sub-primitive type computation graphs can be transformed into at least one corresponding target type computation graph.
[0065] For example, as described above, a first-level subprime type computation graph can be a first-level sub-dynamic computation graph, and a target type computation graph can be a static computation graph. Based on this, the first-level subprime type computation graph is transformed into at least one corresponding target type computation graph, that is, the first-level sub-dynamic computation graph is transformed into at least one corresponding sub-static computation graph. Wherein, if the first-level subprime type computation graph includes sub-dynamic computation... Figure 1 Sub-dynamic computation Figure 2 Dynamic calculation of sums Figure 3 ,like Figure 5 As shown, sub-dynamic computation can be performed. Figure 3 Transform into the corresponding sub-static computation Figure 3 '.
[0066] Step S404: Generate a file to be inferred from the target type computation graph corresponding to the first-level sub-primary type computation graph, and send the file to be inferred to the remote device. The file to be inferred is a file provided to the remote device for inference.
[0067] In this embodiment, after transforming the first-level subprime type computation graph into at least one corresponding target type computation graph, a file to be inferred can be generated from the target type computation graph corresponding to the first-level subprime type computation graph, and the file to be inferred can be sent to a remote device so that the remote device can perform inference operations based on the file to be inferred to obtain the inference result.
[0068] In this embodiment, the file to be inferred can be a model file, wherein the model file can be binary data, and the remote device can be a remote computer terminal.
[0069] For example, after transforming the first-level sub-dynamic computation graph into at least one corresponding sub-static computation graph, a model file can be generated from the at least one sub-static computation graph, and the model file can be sent to a remote device. The remote device can deploy the model file, load the model file, perform inference operations, and obtain the inference results.
[0070] Optionally, when performing remote inference, an inference request can be sent to a remote device. The inference request includes the input data required for inference. After receiving the inference request, the remote device can respond to the inference request, load the model file to perform the inference operation, and obtain the inference result.
[0071] Optionally, after the remote device completes the inference operation and obtains the inference result, the remote device can also send the obtained inference result back to the local device to realize the feedback of the inference result.
[0072] Based on the scheme disclosed in steps S401 to S404 of the above embodiments, by converting the original type computation graph into the corresponding target type computation graph, when the conversion fails, the original type computation graph can be divided into first-level sub-original type computation graphs. Since the first-level sub-original type computation graphs can be further divided into operators defined according to the deep learning framework, and operators defined according to the deep learning framework can necessarily be converted into target type computation graphs, the conversion after dividing the original type computation graph can improve the success rate of the conversion. After converting the first-level sub-original type computation graph into at least one corresponding target type computation graph, the target type computation graph corresponding to the first-level sub-original type computation graph can be used to generate a file to be inferred, and the file to be inferred can be sent to a remote device through the network so that the remote device can perform inference operations based on the file to be inferred and obtain the inference result. This can realize the conversion of any dynamic computation graph into a static computation graph, thereby generating a file to be inferred for inference, without the need for manual modification of the static computation graph. The operation method is relatively simple, achieving the purpose of converting any dynamic computation graph into a static computation graph for inference, thereby achieving the technical effect of improving the inference efficiency of the computation graph, and thus solving the technical problem of low inference efficiency of the computation graph.
[0073] The method described in this embodiment will be further described below.
[0074] As an optional implementation, in step S403, transforming the first-level subprime type computation graph into a corresponding at least one target type computation graph includes: performing at least one-level type conversion on the first-level subprime type computation graph to obtain a corresponding at least one target type computation graph.
[0075] In this embodiment, at least one type transformation is performed on the first-level subprime type computation graph in a recursive manner until at least one corresponding target type computation graph can be successfully obtained.
[0076] In this embodiment, the process of performing at least one-level type conversion on the first-level subprime type computation graph recursively is as follows: First-level type conversion is performed on the first-level subprime type computation graph. If at least one corresponding target type computation graph can be successfully obtained, the conversion stops, and a model file corresponding to the target type computation graph is directly generated. If, after performing first-level type conversion on the first-level subprime type computation graph, the corresponding target type computation graph cannot be obtained, the first-level subprime type computation graph needs to be divided into second-level subprime type computation graphs, and the second-level subprime type computation graph is determined as the new first-level subprime type computation graph. Then, first-level type conversion is performed on the new first-level subprime type computation graph until the corresponding target type computation graph can be successfully obtained.
[0077] For example, such as Figure 5 As shown, dynamic calculation of pairs Figure 3 By performing type conversion, the corresponding sub-static computation can be obtained. Figure 3 ', Pair dynamic calculation Figure 2 If type conversion fails to yield the corresponding sub-static computation graph, then the sub-dynamic computation graph can be used instead. Figure 2 Further division yields sub-dynamic computations. Figure 2 The corresponding sub-dynamic computation graphs 21 and 22 are then converted to obtain sub-static computation graphs 21' and 22'.
[0078] In this embodiment, by performing at least one type transformation on the first-level sub-primary type computation graph to obtain at least one corresponding target type computation graph, the transformation can be attempted multiple times recursively until the sub-dynamic computation graph is successfully converted into the corresponding sub-static computation graph. This allows the sub-static computation graph to generate a file to be inferred, and enables the remote terminal device to perform inference operations based on the file to be inferred to obtain the inference result, without the need for manual modification of the static computation graph, thus improving the efficiency of inference on the computation graph.
[0079] As an optional implementation, performing at least one type conversion on the first-level subprime type computation graph to obtain at least one corresponding target type computation graph includes: a conversion step, performing a type conversion on the first-level subprime type computation graph; a determination step, determining whether the conversion of the first-level subprime type computation graph to the corresponding target type computation graph was successful; a partitioning step, in response to the failure of the conversion of the first-level subprime type computation graph to the corresponding target type computation graph, partitioning the first-level subprime type computation graph into second-level subprime type computation graphs, and determining the second-level subprime type computation graphs as first-level subprime type computation graphs, and returning to the conversion step; and an acquisition step, in response to the successful conversion of the first-level subprime type computation graph to the corresponding target type computation graph, acquiring the target type computation graph corresponding to the first-level subprime type computation graph.
[0080] In this embodiment, the second-level subprime type computation graph is a subgraph of the first-level subprime type computation graph. The first-level subprime type computation graph is transformed, and the transformation is determined to be successful. If the transformation fails, the first-level subprime type computation graph is divided into second-level subprime type computation graphs, and the second-level subprime type computation graphs are identified as first-level subprime type computation graphs. The process returns to the transformation step, and the newly identified first-level subprime type computation graph is transformed. If the transformation fails again, the newly identified first-level subprime type computation graph is again divided into second-level subprime type computation graphs, and the second-level subprime type computation graphs are identified as first-level subprime type computation graphs. The process continues to return to the transformation step. Based on this, a recursive approach is used to perform at least one-level type transformation on the first-level subprime type computation graph until the first-level subprime type computation graph is successfully transformed into the corresponding target type computation graph, thereby obtaining the target type computation graph corresponding to the first-level subprime type computation graph.
[0081] For example, such as Figure 5 As shown, dynamic calculation of pairs Figure 2 Perform the transformation and determine the sub-dynamic calculation. Figure 2 Whether the conversion to the corresponding sub-static computation graph was successful; if the conversion fails, the sub-dynamic computation graph can be... Figure 2The system divides the system into sub-dynamic computation graphs 21 and 22, and identifies them as first-level sub-dynamic computation graphs. The system then returns to the transformation step to transform sub-dynamic computation graphs 21 and 22. If sub-dynamic computation graph 21 is successfully transformed into its corresponding sub-static computation graph 21', but the transformation of sub-dynamic computation graph 22 into its corresponding sub-static computation graph fails, the system further divides the system into sub-dynamic computation graphs 221 and 222, and identifies them as first-level sub-dynamic computation graphs. The system continues to return to the transformation step to transform sub-dynamic computation graphs 221 and 222 until sub-dynamic computation graph 221 is successfully transformed into its corresponding sub-static computation graph 221' and sub-dynamic computation graph 222'.
[0082] In this embodiment, a recursive approach is used to perform at least one level of type conversion on the first-level subprime type computation graph until the first-level subprime type computation graph is successfully converted into the corresponding target type computation graph. This allows the target type computation graph corresponding to the first-level subprime type computation graph to be obtained. The conversion can be attempted multiple times through recursion until the sub-dynamic computation graph is successfully converted into the sub-static computation graph. This avoids the problem that the dynamic computation graph cannot be converted into the static computation graph, and eliminates the need for manual modification of the static computation graph, thereby improving the efficiency of reasoning on the computation graph.
[0083] As an optional implementation, the original type computation graph is obtained by the target operator, which is used to successfully convert to the corresponding target type computation graph. After dividing the first-level sub-original type computation graph into second-level sub-original type computation graphs, the computation graph reasoning method further includes: in response to the second-level sub-original type computation graph being the target operator, obtaining the target type computation graph successfully converted by the target operator, and stopping the execution of the following step: determining the second-level sub-original type computation graph as the first-level sub-original type computation graph.
[0084] In this embodiment, the target operator is the basic unit for building a primitive type computation graph. It can be the basic unit for building a dynamic computation graph and can include operators defined by the deep learning framework and basic operations. The basic operations can include addition, subtraction, multiplication, and division. Based on this, when a sub-primitive type computation graph is assigned to a target operator, the target operator can be successfully converted into a static computation graph. Therefore, when a second-level sub-primitive type computation graph is a target operator, since the target operator can be converted into a target type computation graph, the target type computation graph successfully converted by the target operator can be directly obtained without further determining the second-level sub-primitive type computation graph as a first-level sub-primitive type computation graph.
[0085] For example, such as Figure 5As shown, sub-dynamic computation Figure 1 When it cannot be converted into the corresponding sub-static computation graph, the sub-dynamic computation graph can be used. Figure 1 The system is divided into sub-dynamic computation graph 11 and sub-dynamic computation graph 12. When sub-dynamic computation graph 11 and sub-dynamic computation graph 12 are operators defined by the deep learning framework, the sub-static computation graph 11' and sub-static computation graph 12' successfully converted from sub-dynamic computation graph 11 and sub-static computation graph 12 can be obtained directly without further determining sub-dynamic computation graph 11 and sub-dynamic computation graph 12 as first-level sub-dynamic computation graphs.
[0086] In this embodiment, by dividing the sub-dynamic computation graph multiple times, a multi-level sub-dynamic computation graph can be obtained. When the sub-dynamic computation graph is divided to the innermost layer, that is, when the sub-dynamic computation graph is divided into multiple operators defined by deep learning frameworks, the multiple operators defined by deep learning frameworks can be directly converted into the corresponding sub-static computation graphs. This enables any dynamic computation graph to be converted into a static computation graph, thereby avoiding the problem that dynamic computation graphs cannot be converted into static computation graphs.
[0087] As an optional implementation, the reasoning method for the computation graph further includes: constructing a target operator based on a deep learning framework, wherein the deep learning framework is used to convert the target operator into a corresponding target type computation graph.
[0088] In this embodiment, a target operator can be constructed based on a deep learning framework, wherein the deep learning framework can be used to convert the target operator into a corresponding target type computation graph.
[0089] For example, when converting a target operator into its corresponding sub-static computation graph, a deep learning framework can be used for the conversion operation.
[0090] As an optional implementation, in step S403, the first-level subprime type computation graph is transformed into a corresponding at least one target type computation graph, including: transforming the first-level subprime type computation graph into a corresponding at least one target type computation graph based on the interface provided by the deep learning framework.
[0091] In this embodiment, the first-level subprime type computation graph can be transformed into a corresponding at least one target type computation graph based on the interface provided by the deep learning framework.
[0092] For example, when converting a sub-dynamic computation graph into a sub-static computation graph, the conversion can be achieved using the interfaces provided by the deep learning framework.
[0093] As an optional implementation, the original type computation graph can be converted into the corresponding target type computation graph based on the interface provided by the deep learning framework.
[0094] In this embodiment, the original type computation graph can also be converted into the corresponding target type computation graph based on the interface provided by the deep learning framework.
[0095] For example, you can use the interfaces provided by deep learning frameworks to convert dynamic computation graphs into corresponding static computation graphs.
[0096] As an optional implementation, based on the anomaly alerts output by the deep learning framework, it is determined that the original type computation graph failed to be converted into the corresponding target type computation graph.
[0097] In this embodiment, it can also be determined, based on the abnormal prompt information output by the deep learning framework, that the detected failure to convert the original type computation graph into the corresponding target type computation graph has been identified.
[0098] For example, when the conversion of a dynamic computation graph to a static computation graph fails using the interface provided by the deep learning framework, the interface of the deep learning framework will automatically output an exception message, which can then be used to determine whether the conversion of the dynamic computation graph to a static computation graph has failed.
[0099] As an optional implementation, in step S402, the primitive type computation graph is divided into first-level sub-primitive type computation graphs, which includes: dividing the primitive type computation graph into first-level sub-primitive type computation graphs based on the format corresponding to the deep learning framework, wherein the primitive type computation graph and the first-level sub-primitive type computation graphs satisfy the format corresponding to the deep learning framework.
[0100] In this embodiment, the primitive type computation graph can be divided into first-level sub-primitive type computation graphs based on the format corresponding to the deep learning framework, wherein the primitive type computation graph and the first-level sub-primitive type computation graph satisfy the format corresponding to the deep learning framework.
[0101] For example, when a dynamic computation graph conforms to the format corresponding to a deep learning framework, it can be divided into first-level sub-dynamic computation graphs according to the format of the deep learning framework. The first-level sub-dynamic computation graphs also conform to the format of the deep learning framework.
[0102] As an optional implementation, in response to the successful conversion of the original type computation graph into the corresponding target type computation graph, a file to be inferred is generated from the converted target type computation graph.
[0103] In this embodiment, when the detected original type computation graph is successfully converted into the corresponding target type computation graph, the converted target type computation graph can be used to generate a file to be inferred.
[0104] For example, when a dynamic computation graph is successfully converted into a corresponding static computation graph, a model file can be generated from the static computation graph. The model file can then be sent to a remote computer over a network, so that the remote computer can perform inference operations based on the model file and obtain the inference results.
[0105] In the above steps, by converting the original type computation graph into the corresponding target type computation graph, when the conversion fails, the original type computation graph can be divided into first-level sub-original type computation graphs. Since the first-level sub-original type computation graphs can be further divided into operators defined according to the deep learning framework, and operators defined according to the deep learning framework can necessarily be converted into target type computation graphs, dividing the original type computation graph before conversion can improve the success rate of conversion. After converting the first-level sub-original type computation graph into at least one corresponding target type computation graph, the target type computation graph corresponding to the first-level sub-original type computation graph can be used to generate a file to be inferred. The file to be inferred is then sent to a remote device via the network, so that the remote device can perform inference operations based on the file to be inferred and obtain the inference result. This can realize the conversion of any dynamic computation graph into a static computation graph, thereby generating a file to be inferred for inference, without the need for manual modification of the static computation graph. The operation method is relatively simple, achieving the purpose of converting any dynamic computation graph into a static computation graph for inference, thereby achieving the technical effect of improving the inference efficiency of computation graphs, and thus solving the technical problem of low inference efficiency of computation graphs.
[0106] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0107] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0108] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0109] Example 2
[0110] The preferred embodiments of the method described above in this example will be further described below.
[0111] Currently, when using deep learning frameworks for remote inference, it is typically necessary to convert dynamic computation graphs into static computation graphs, export model files, and finally perform remote inference. However, in many scenarios, dynamic computation graphs cannot be successfully converted into static computation graphs. For example, if the dynamic computation graph contains Python-specific syntax, or if the input and output data of the dynamic computation graph are not types supported by the deep learning framework, the conversion from dynamic to static computation graph will fail.
[0112] In related technologies, when a dynamic computation graph cannot be converted into a static computation graph, the usual solution is for the user to manually modify the static computation graph to ensure a successful conversion. Therefore, when using deep learning frameworks for remote inference, in many scenarios, it is impossible to directly convert a dynamic computation graph into a static one, which easily leads to the technical problem of low efficiency in remote inference.
[0113] To address the aforementioned issues, this application proposes a computation graph inference method. This method involves converting a primitive type computation graph into a corresponding target type computation graph. If the conversion fails, the primitive type computation graph can be divided into first-level sub-primitive computation graphs. Since these first-level sub-primitive type computation graphs can be further divided into operators defined by deep learning frameworks, and these operators can be converted into target type computation graphs, dividing the primitive type computation graph before conversion improves the success rate. After converting the first-level sub-primitive type computation graph into at least one corresponding target type computation graph, a file to be inferred can be generated from the target type computation graph corresponding to the first-level sub-primitive type computation graph. This file is then sent to a remote device via a network, allowing the remote device to perform inference operations based on the file and obtain the inference result. This method enables the conversion of any dynamic computation graph into a static computation graph, generating a file to be inferred for inference without requiring manual modification of the static computation graph. The operation is relatively simple, achieving the goal of converting any dynamic computation graph into a static computation graph for inference. This improves the inference efficiency of computation graphs and solves the problem of low inference efficiency.
[0114] The reasoning behind computational graphs will be further explained below.
[0115] Figure 6 This is a flowchart of another computational graph reasoning method according to an embodiment of this application. For example... Figure 6As shown, the process first monitors the primitive type computation graph to be inferred and converts it into the corresponding target type computation graph using the interface provided by the deep learning framework. Then, it determines whether the conversion was successful. If successful, the target type computation graph is directly generated as the inference file. If conversion fails, the primitive type computation graph is divided into first-level sub-primitive type computation graphs based on the format of the deep learning framework. These first-level sub-primitive type computation graphs also conform to the deep learning framework's format. Subsequently, using the interface provided by the deep learning framework, all first-level sub-primitive type computation graphs are converted into the corresponding target type computation graphs, and the conversion is then determined. If the graph is successfully converted to the corresponding target type computation graph, the first-level subprime type computation graph is directly converted to the corresponding target type computation graph, and the target type computation graph is directly generated into a file to be inferred. If the conversion of the first-level subprime type computation graph fails, all the first-level subprime type computation graphs can be further divided into second-level subprime type computation graphs based on the format corresponding to the deep learning framework. The second-level subprime type computation graphs are then identified as first-level subprime type computation graphs, and the conversion step is returned. That is, the first-level subprime type computation graphs are converted into the corresponding target type computation graphs. The first-level subprime type computation graphs can be divided into second-level subprime type computation graphs in a recursive manner until the second-level subprime type computation graphs can be converted into the corresponding target type computation graphs. Finally, the file to be inferred can be generated.
[0116] Figure 5 This is a schematic diagram of a computational graph reasoning method according to an embodiment of this application. For example... Figure 5 As shown, when the detected dynamic computation graph cannot be directly converted into a static computation graph, the dynamic computation graph can be divided into sub-dynamic computation graphs based on the format corresponding to the deep learning framework. Figure 1 Sub-dynamic computation Figure 2 and sub-dynamic computation Figure 3 Among them, sub-dynamic calculation Figure 3 It can be directly converted into the corresponding sub-static computation. Figure 3 ', and sub-dynamic computation Figure 1 Dynamic calculation of sums Figure 2 It cannot be directly converted into the corresponding sub-static computation graph, therefore the sub-dynamic computation graph will be used instead. Figure 1 The system is divided into sub-dynamic computation graph 11 and sub-dynamic computation graph 12. Sub-dynamic computation graph 11 and sub-dynamic computation graph 12 can be directly converted into corresponding sub-static computation graphs 11' and 12'. Subsequently, the sub-dynamic computation... Figure 2The system is divided into sub-dynamic computation graphs 21 and 22. Sub-dynamic computation graph 21 can be directly converted into its corresponding sub-static computation graph 21', while sub-dynamic computation graph 22 cannot be directly converted into its corresponding sub-static computation graph. Therefore, sub-dynamic computation graph 22 is further divided into sub-dynamic computation graphs 221 and 222. Sub-dynamic computation graphs 221 and 222 can be directly converted into their corresponding sub-static computation graphs 221' and 222'. Finally, through recursion, the dynamic computation graph can be divided into six sub-dynamic computation graphs layer by layer. These six sub-dynamic computation graphs can then be converted into six corresponding sub-static computation graphs. These six converted sub-static computation graphs can then generate six model files, which are sent to a remote computer. The remote computer can then traverse the model files sequentially and perform inference to obtain the inference results.
[0117] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0118] Example 3
[0119] According to embodiments of this application, a computation graph inference system is also provided, such as... Figure 7 As shown, the computation graph inference system includes a local device 701 for monitoring the original type computation graph to be inferred, wherein the original type computation graph is a computation graph constructed and computed within the same time period; in response to the failure of the detected original type computation graph to be converted into the corresponding target type computation graph, the original type computation graph is divided into first-level sub-original type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; the first-level sub-original type computation graphs are converted into at least one corresponding target type computation graph; the target type computation graph corresponding to the first-level sub-original type computation graph is used to generate a file to be inferred; and a remote device 702 is used to obtain the file to be inferred and to perform inference on the file to be inferred to obtain the inference result.
[0120] In this embodiment, the original type computation graph to be inferred can be monitored using a local device. The original type computation graph is a computation graph that is constructed and computed within the same time period. For example, the original type computation graph can be a dynamic computation graph, which can be called a dynamic graph.
[0121] Optionally, after the local device detects the computation graph of the original type to be inferred, the original type computation graph can be directly converted into a target type computation graph using the API provided by the deep learning framework. The target type computation graph can be a static computation graph.
[0122] In this embodiment, after the local device detects the primitive type computation graph to be inferred, it can perform type conversion on the primitive type computation graph detected by the local device. For example, it can perform type conversion from a dynamic computation graph to a static computation graph. In response to the failure of converting the primitive type computation graph detected by the local device into the corresponding target type computation graph, the primitive type computation graph can be divided into first-level sub-primitive type computation graphs.
[0123] In this embodiment, since there may be situations in many scenarios where the primitive type computation graph cannot be directly converted to the target type computation graph—for example, the primitive type computation graph contains syntax specific to Python, a computer programming language that allows for simple and efficient object-oriented programming, or the input and output of the primitive type computation graph are not types supported by the deep learning framework—the API provided by the deep learning framework will automatically throw an exception during the conversion process. When an exception is detected from the API provided by the deep learning framework, i.e., in response to a failure to convert the primitive type computation graph to the corresponding target type computation graph, the primitive type computation graph can then be divided into first-level sub-primitive type computation graphs for conversion.
[0124] Optionally, the aforementioned target type computation graph is a computation graph constructed and computed within different time periods, and can be a static computation graph, also referred to as a static graph. Since the original type computation graph can be a dynamic computation graph, this first-level sub-original type computation graph can be a first-level sub-dynamic computation graph obtained by partitioning the dynamic computation graph.
[0125] Optionally, if the primitive type computation graph conforms to the standard format of a deep learning framework, then the type and number of the first-level sub-primitive type computation graphs obtained after partitioning the primitive type computation graph are determined. For example, if the primitive type computation graph conforms to the standard format of a deep learning framework, then the first-level sub-primitive type computation graphs obtained after partitioning the primitive type computation graph can be as follows: Figure 5 Sub-dynamic calculation shown Figure 1 Sub-dynamic computation Figure 2 and sub-dynamic computation Figure 3 The examples provided are merely illustrative and do not constitute a limitation on the embodiments of this application.
[0126] In this embodiment, a first-level sub-primary type computation graph can be transformed into a corresponding at least one target type computation graph in the local device. For example, a sub-dynamic computation graph can be transformed into a corresponding at least one sub-static computation graph in the local device.
[0127] For example, as described above, a first-level subprime type computation graph can be a first-level sub-dynamic computation graph, and a target type computation graph can be a static computation graph. Based on this, the first-level subprime type computation graph is transformed into at least one corresponding target type computation graph, that is, the first-level sub-dynamic computation graph is transformed into at least one corresponding sub-static computation graph. Wherein, if the first-level subprime type computation graph includes sub-dynamic computation... Figure 1 Sub-dynamic computation Figure 2 Dynamic calculation of sums Figure 3 ,like Figure 5 As shown, sub-dynamic computation can be performed. Figure 3 Transform into the corresponding sub-static computation Figure 3 '.
[0128] In this embodiment, after the local device transforms the first-level subprime type computation graph into at least one corresponding target type computation graph, it can generate a file to be inferred from the target type computation graph corresponding to the first-level subprime type computation graph. The file to be inferred can be a model file, for example, binary data.
[0129] In this embodiment, after the local device generates a file to be inferred from the target type computation graph corresponding to the first-level sub-primary type computation graph, the remote device can obtain the file to be inferred via the network and perform inference on the file to obtain the inference result.
[0130] Optionally, the local device 701 is also used to send the file to be inferred to a remote device based on a deep learning framework.
[0131] In this embodiment, the local device 701 can also send the file to be inferred to a remote device based on a deep learning framework, so that the remote device can load the file to be inferred to perform inference operations and obtain inference results.
[0132] Optionally, the local device 701 is also used to send an inference request to a remote device based on a deep learning framework, and the remote device responds to the inference request, performs inference on the inference file, and obtains the inference result.
[0133] In this embodiment, the local device 701 can also send an inference request to the remote device based on a deep learning framework. After receiving the inference request, the remote device can respond to the inference request, perform inference operations on the inference file, and obtain the inference result.
[0134] Based on the above embodiments, by converting the original type computation graph into a corresponding target type computation graph in the local device, when the conversion fails, the original type computation graph can be divided into first-level sub-original type computation graphs. Since the first-level sub-original type computation graphs can be further divided into operators defined according to the deep learning framework, and operators defined according to the deep learning framework can necessarily be converted into target type computation graphs, the success rate of conversion can be improved after dividing the original type computation graph into first-level sub-original type computation graphs. After converting the first-level sub-original type computation graphs into at least one corresponding target type computation graph, a file to be inferred can be generated from the target type computation graph corresponding to the first-level sub-original type computation graph. The file to be inferred is then sent to a remote device via the network, so that the remote device can perform inference operations based on the file to be inferred and obtain the inference result. This can realize the conversion of any dynamic computation graph into a static computation graph, thereby generating a file to be inferred for inference, without the need for manual modification of the static computation graph. The operation method is relatively simple, achieving the purpose of converting any dynamic computation graph into a static computation graph for inference, thereby achieving the technical effect of improving the inference efficiency of computation graphs, and thus solving the technical problem of low inference efficiency of computation graphs.
[0135] Example 4
[0136] According to embodiments of this application, a method for implementing the above is also provided. Figure 4 The inference device of the computational graph is shown as a reasoning method for the computational graph.
[0137] Figure 8 This is a structural block diagram of a computational graph inference device according to an embodiment of this application. For example... Figure 8 As shown, the inference device of the computation graph may include: a monitoring module 801, a partitioning module 802, a conversion module 803, and a sending module 804.
[0138] The monitoring module 801 is used to monitor the original type computation graph to be inferred, wherein the original type computation graph is a computation graph that is constructed and computed within the same time period.
[0139] The partitioning module 802 is used to partition the original type computation graph into first-level sub-original type computation graphs in response to the detected failure to convert the original type computation graph into the corresponding target type computation graph. The target type computation graph is a computation graph that is constructed and computed in different time periods.
[0140] The transformation module 803 is used to transform the first-level sub-primary type computation graph into a corresponding target type computation graph.
[0141] The sending module 804 is used to generate a file to be inferred from the target type computation graph corresponding to the first-level sub-primary type computation graph, and send the file to be inferred to a remote device, wherein the file to be inferred is a file provided to the remote device for inference.
[0142] Optionally, the conversion module 803 is further configured to convert the first-level subprime type computation graph into a corresponding at least one target type computation graph, including: performing at least one-level type conversion on the first-level subprime type computation graph to obtain a corresponding at least one target type computation graph.
[0143] Optionally, the conversion module 803 is further configured to perform at least one type conversion on the first-level subprime type computation graph to obtain at least one corresponding target type computation graph, including: a conversion step, performing type conversion on the first-level subprime type computation graph; a determination step, determining whether the conversion of the first-level subprime type computation graph to the corresponding target type computation graph is successful; a partitioning step, in response to the failure of the conversion of the first-level subprime type computation graph to the corresponding target type computation graph, partitioning the first-level subprime type computation graph into a second-level subprime type computation graph, and determining the second-level subprime type computation graph as a first-level subprime type computation graph, and returning to the conversion step; and an acquisition step, in response to the successful conversion of the first-level subprime type computation graph to the corresponding target type computation graph, acquiring the target type computation graph corresponding to the first-level subprime type computation graph.
[0144] Optionally, the conversion module 803 is further configured to, in response to the second-level subprime type computation graph being the target operator, obtain the target type computation graph successfully converted from the target operator, and stop executing the following steps: determining the second-level subprime type computation graph as the first-level subprime type computation graph.
[0145] Optionally, the transformation module 803 is also used to construct a target operator based on a deep learning framework, wherein the deep learning framework is used to convert the target operator into a corresponding target type computation graph.
[0146] Optionally, the transformation module 803 is further configured to transform the first-level subprime type computation graph into a corresponding at least one target type computation graph, including: transforming the first-level subprime type computation graph into a corresponding at least one target type computation graph based on the interface provided by the deep learning framework.
[0147] Optionally, the conversion module 803 is also used to convert the original type computation graph into the corresponding target type computation graph based on the interface provided by the deep learning framework.
[0148] Optionally, the inference device for the computation graph further includes a determination module 805, which determines, based on the abnormal prompt information output by the deep learning framework, that the detected original type computation graph failed to be converted into the corresponding target type computation graph.
[0149] Optionally, the partitioning module 802 is further configured to partition the primitive type computation graph into first-level sub-primitive type computation graphs, including: partitioning the primitive type computation graph into first-level sub-primitive type computation graphs based on the format corresponding to the deep learning framework, wherein the primitive type computation graph and the first-level sub-primitive type computation graph satisfy the format corresponding to the deep learning framework.
[0150] Optionally, the inference apparatus for the computation graph further includes a generation module 806, which generates a file to be inferred from the converted target type computation graph in response to the successful conversion of the original type computation graph into the corresponding target type computation graph.
[0151] It should be noted that the monitoring module 801, the division module 802, the conversion module 803 and the sending module 804 mentioned above correspond to steps S401 to S404 in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and can run in the device provided in Embodiment 1.
[0152] Example 5
[0153] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.
[0154] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0155] In this embodiment, the computer terminal described above can execute the program code for the following steps in the computation graph inference method: monitoring the primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to the failure of the monitored primitive type computation graph to be converted into the corresponding target type computation graph, dividing the primitive type computation graph into first-level sub-primary type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; converting the first-level sub-primary type computation graphs into at least one corresponding target type computation graph; generating a file to be inferred from the target type computation graph corresponding to the first-level sub-primary type computation graph, and sending the file to be inferred to a remote device, wherein the file to be inferred is a file provided to the remote device for inference.
[0156] Optionally, Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 9As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 902, memory 904, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.
[0157] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the computation graph inference method in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned computation graph inference method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0158] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: monitoring the primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to the failure of the detected primitive type computation graph to be converted into the corresponding target type computation graph, dividing the primitive type computation graph into first-level sub-primitive type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; converting the first-level sub-primitive type computation graphs into at least one corresponding target type computation graph; generating a file to be inferred from the target type computation graph corresponding to the first-level sub-primitive type computation graph, and sending the file to be inferred to a remote device via a network, wherein the file to be inferred is a file provided to the remote device for inference.
[0159] Optionally, the processor may also execute program code that performs at least one type conversion on the first-level subprime type computation graph to obtain at least one corresponding target type computation graph.
[0160] Optionally, the processor may also execute program code with the following steps: a conversion step, performing a type conversion on the first-level subprime type computation graph; a determination step, determining whether the conversion of the first-level subprime type computation graph to the corresponding target type computation graph was successful; a partitioning step, in response to the failure of the conversion of the first-level subprime type computation graph to the corresponding target type computation graph, partitioning the first-level subprime type computation graph into a second-level subprime type computation graph, and determining the second-level subprime type computation graph as a first-level subprime type computation graph, and returning to the conversion step; and an acquisition step, in response to the successful conversion of the first-level subprime type computation graph to the corresponding target type computation graph, acquiring the target type computation graph corresponding to the first-level subprime type computation graph.
[0161] Optionally, the processor may also execute program code that performs the following steps: in response to the second-level subprime type computation graph being the target operator, obtain the target type computation graph successfully converted by the target operator, and stop performing the following steps: determine the second-level subprime type computation graph as the first-level subprime type computation graph.
[0162] Optionally, the processor may also execute program code that performs the following steps: constructing a target operator based on a deep learning framework, wherein the deep learning framework is used to convert the target operator into a corresponding target type computation graph.
[0163] Optionally, the processor may also execute program code that performs the following steps: transforms a first-level subprime type computation graph into a corresponding at least one target type computation graph based on the interface provided by the deep learning framework.
[0164] Optionally, the processor may also execute program code that converts the original type computation graph into the corresponding target type computation graph based on the interface provided by the deep learning framework.
[0165] Optionally, the processor may also execute program code that performs the following steps: based on the exception message output by the deep learning framework, determine that the detected failure to convert the original type computation graph into the corresponding target type computation graph.
[0166] Optionally, the processor may also execute program code that performs the following steps: based on the format corresponding to the deep learning framework, the primitive type computation graph is divided into first-level sub-primitive type computation graphs, wherein the primitive type computation graph and the first-level sub-primitive type computation graph satisfy the format corresponding to the deep learning framework.
[0167] Optionally, the processor may also execute program code that performs the following steps: in response to the successful conversion of the detected original type computation graph into the corresponding target type computation graph, generates a file to be inferred from the converted target type computation graph.
[0168] This application provides an inference scheme for computation graphs. By converting a primitive type computation graph into a corresponding target type computation graph, when conversion fails, the primitive type computation graph can be divided into first-level sub-primitive computation graphs. Since these first-level sub-primitive type computation graphs can be further divided into operators defined by deep learning frameworks, and operators defined by deep learning frameworks can necessarily be converted into target type computation graphs, dividing the primitive type computation graph before conversion improves the success rate. After converting the first-level sub-primitive type computation graph into at least one corresponding target type computation graph, a file to be inferred can be generated from the target type computation graph corresponding to the first-level sub-primitive type computation graph. This file is then sent to a remote device via a network, allowing the remote device to perform inference and obtain the inference result. This scheme enables the conversion of any dynamic computation graph into a static graph, thereby generating a file to be inferred for inference without manual modification of the static computation graph. The operation is relatively simple, achieving the goal of converting any dynamic computation graph into a static computation graph for inference, thus improving the inference efficiency of computation graphs and solving the technical problem of low inference efficiency.
[0169] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an Apple operating system (iOS) phone, a tablet computer, a PDA, a mobile internet device (MID), a laptop computer, or other terminal devices. Figure 9 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more... Figure 9 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 9 The different configurations shown.
[0170] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0171] Example 6
[0172] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the aforementioned computer-readable storage medium can be used to store the program code executed by the inference method of the computation graph provided in Embodiment 1 above.
[0173] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0174] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: monitoring the primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to the failure of the monitored primitive type computation graph to be converted into the corresponding target type computation graph, dividing the primitive type computation graph into first-level sub-primitive type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; converting the first-level sub-primitive type computation graphs into at least one corresponding target type computation graph; generating a file to be inferred from the target type computation graph corresponding to the first-level sub-primitive type computation graph, and sending the file to be inferred to a remote device, wherein the file to be inferred is a file provided to the remote device for inference.
[0175] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing at least one level type conversion on the first-level subprime type computation graph to obtain at least one corresponding target type computation graph.
[0176] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: a conversion step, performing a type conversion on a first-level subprime type computation graph; a determination step, determining whether the conversion of the first-level subprime type computation graph to the corresponding target type computation graph was successful; a partitioning step, in response to the failure of the conversion of the first-level subprime type computation graph to the corresponding target type computation graph, partitioning the first-level subprime type computation graph into a second-level subprime type computation graph, and determining the second-level subprime type computation graph as a first-level subprime type computation graph, and returning to the conversion step; and an acquisition step, in response to the successful conversion of the first-level subprime type computation graph to the corresponding target type computation graph, acquiring the target type computation graph corresponding to the first-level subprime type computation graph.
[0177] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in response to the second-level subprime type computation graph being the target operator, obtaining the successful conversion of the target operator into the corresponding target type computation graph, and stopping the execution of the following steps: determining the second-level subprime type computation graph as the first-level subprime type computation graph.
[0178] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: constructing a target operator based on a deep learning framework, wherein the deep learning framework is used to convert the target operator into a corresponding target type computation graph.
[0179] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: transforming a first-level subprime type computation graph into a corresponding at least one target type computation graph based on an interface provided by a deep learning framework.
[0180] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: converting the original type computation graph into a corresponding target type computation graph based on the interface provided by the deep learning framework.
[0181] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining, based on the exception prompt information output by the deep learning framework, that the detected original type computation graph failed to be converted into the corresponding target type computation graph.
[0182] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: dividing the primitive type computation graph into first-level sub-primitive type computation graphs based on the format corresponding to the deep learning framework, wherein the primitive type computation graph and the first-level sub-primitive type computation graph satisfy the format corresponding to the deep learning framework.
[0183] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in response to the successful conversion of the detected original type computation graph into the corresponding target type computation graph, generating a file to be inferred from the converted target type computation graph.
[0184] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0185] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0186] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0188] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0189] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0190] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A reasoning method for computational graphs, characterized in that, include: Monitor the primitive type computation graph to be reasoned, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; In response to the detected failure to convert the original type computation graph into the corresponding target type computation graph, the original type computation graph is divided into first-level sub-original type computation graphs, wherein the target type computation graph is a computation graph constructed and computed in different time periods; Transform the first-level subprime type computation graph into at least one corresponding target type computation graph; The target type computation graph corresponding to the first-level sub-primitive type computation graph is used to generate a file to be inferred, and the file to be inferred is sent to a remote device, wherein the file to be inferred is a file provided to the remote device for inference; The first-level subprime type computation graph is a first-level sub-dynamic computation graph, and the first-level subprime type computation graph is used to partition into operators. The operators are defined according to the deep learning framework and are used to transform the target type computation graph.
2. The method according to claim 1, characterized in that, Transforming the first-level subprime type computation graph into at least one corresponding target type computation graph includes: Perform at least one type transformation on the first-level sub-primitive type computation graph to obtain at least one corresponding target type computation graph.
3. The method according to claim 2, characterized in that, Perform at least one type transformation on the first-level subprime type computation graph to obtain at least one corresponding target type computation graph, including: The conversion step involves performing a type conversion on the first-level sub-primary type computation graph; The steps include determining whether the conversion of the first-level sub-primary type computation graph into the corresponding target type computation graph was successful. In the partitioning step, in response to the failure of converting the first-level subprime type computation graph into the corresponding target type computation graph, the first-level subprime type computation graph is partitioned into a second-level subprime type computation graph, and the second-level subprime type computation graph is identified as the first-level subprime type computation graph, and the process returns to the conversion step. In the acquisition step, in response to the successful conversion of the first-level subprime type computation graph into the corresponding target type computation graph, the target type computation graph corresponding to the first-level subprime type computation graph is acquired.
4. The method according to claim 3, characterized in that, The primitive type computation graph is obtained by the target operator, which is used to successfully convert it into the corresponding target type computation graph. The method further includes, after dividing the first-level sub-primitive type computation graph into second-level sub-primitive type computation graphs: In response to the second-level subprime type computation graph being the target operator, the target type computation graph successfully converted by the target operator is obtained, and the following steps are stopped: the second-level subprime type computation graph is determined as the first-level subprime type computation graph.
5. The method according to claim 4, characterized in that, The method further includes: The target operator is constructed based on a deep learning framework, wherein the deep learning framework is used to convert the target operator into a corresponding target type computation graph.
6. The method according to claim 1, characterized in that, Transforming the first-level subprime type computation graph into at least one corresponding target type computation graph includes: Based on the interface provided by the deep learning framework, the first-level subprime type computation graph is transformed into at least one corresponding target type computation graph.
7. The method according to claim 1, characterized in that, The method further includes: Based on the interface provided by the deep learning framework, the original type computation graph is converted into the corresponding target type computation graph.
8. The method according to claim 1, characterized in that, The method further includes: Based on the abnormal prompts output by the deep learning framework, it is determined that the original type computation graph failed to be converted into the corresponding target type computation graph.
9. The method according to claim 1, characterized in that, The primitive type computation graph is divided into first-level sub-primitive type computation graphs, including: Based on the format corresponding to the deep learning framework, the primitive type computation graph is divided into the first-level sub-primitive type computation graph, wherein the primitive type computation graph and the first-level sub-primitive type computation graph satisfy the format corresponding to the deep learning framework.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: In response to the successful conversion of the original type computation graph into the corresponding target type computation graph, the converted target type computation graph is used to generate the inference file.
11. A computational graph reasoning system, characterized in that, include: A local device is used to monitor the primitive type computation graph to be inferred, wherein the primitive type computation graph is a computation graph constructed and computed within the same time period; in response to the detected failure to convert the primitive type computation graph into a corresponding target type computation graph, the primitive type computation graph is divided into first-level sub-primitive type computation graphs, wherein the target type computation graphs are computation graphs constructed and computed in different time periods; the first-level sub-primitive type computation graphs are converted into at least one corresponding target type computation graph; and the inference file generated from the target type computation graph corresponding to the first-level sub-primitive type computation graph is generated. A remote device is used to acquire the file to be inferred, and to perform inference on the file to be inferred to obtain the inference result; The first-level subprime type computation graph is a first-level sub-dynamic computation graph, and the first-level subprime type computation graph is used to partition into operators. The operators are defined according to the deep learning framework and are used to transform the target type computation graph.
12. The system according to claim 11, characterized in that, The local device is used to send the file to be inferred to the remote device based on a deep learning framework.
13. The system according to claim 11, characterized in that, The local device is used to send an inference request to the remote device based on a deep learning framework, and the remote device is used to respond to the inference request, perform inference on the inference file, and obtain the inference result.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program is run by a processor, it controls the device in which the computer-readable storage medium resides to perform the method according to any one of claims 1 to 10.