Model processing method, device, electronic device and storage medium

By classifying and sub-graphing the kernel functions of the deep learning model, the problem of low optimization efficiency caused by mismatch between the kernel functions and the optimization algorithm is solved, and the optimization efficiency of the model is improved.

CN115293228BActive Publication Date: 2025-08-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210706927.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-08-26
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

In the prior art, the optimization algorithm of the deep learning model leads to inefficient optimization efficiency when some kernel functions do not match the optimization algorithm.

Method used

By obtaining the calculation order of the kernel functions in the initial model and the target algorithm, classifying the kernel functions, determining the target subgraph, and optimizing the initial model using the target algorithm and the target subgraph to remove mismatched kernel functions.

Benefits of technology

The kernel functions matching the optimization algorithm in the model are optimized, removing mismatched kernel functions and improving optimization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293228B_ABST
    Figure CN115293228B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model processing method, apparatus, electronic device, and storage medium, relating to the fields of artificial intelligence technology and, more specifically, deep learning, to at least address the low optimization efficiency of models in related technologies. The specific implementation scheme comprises: obtaining the calculation order and target algorithm of multiple kernel functions in an initial model; classifying the multiple kernel functions to obtain classification results; determining a target subgraph based on the classification results and the calculation order; and optimizing the initial model using the target algorithm and target subgraph to obtain a target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a model processing method, device, electronic device, and storage medium. Background Art

[0002] The deep learning models in the existing technology all optimize the entire model. However, when some kernel functions in the model do not match the optimization algorithm, the optimization algorithm will still optimize the mismatched kernel functions, which will lead to low optimization efficiency of the model optimization algorithm. Summary of the Invention

[0003] The present disclosure provides a model processing method, device, electronic device and storage medium to at least solve the technical problem of low optimization efficiency of model optimization in related technologies.

[0004] According to one aspect of the present disclosure, a model processing method is provided, including: obtaining a calculation order and a target algorithm of multiple kernel functions in an initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model; classifying the multiple kernel functions to obtain a classification result, wherein the classification result is used to indicate whether the multiple kernel functions support the target algorithm; determining a target subgraph based on the classification result and the calculation order, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function that supports the target algorithm; optimizing the initial model using the target algorithm and the target subgraph to obtain a target model.

[0005] According to another aspect of the present disclosure, a model processing device is provided, including: an acquisition module for acquiring the calculation order and target algorithm of multiple kernel functions in an initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model; a classification module for classifying the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm; a determination module for determining a target subgraph based on the classification results and the calculation order, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function that supports the target algorithm; an optimization module for optimizing the initial model using the target algorithm and the target subgraph to obtain the target model.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model processing method proposed in the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the model processing method proposed in the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which executes the model processing method proposed in the present disclosure when a processor executes the computer program.

[0009] In the present disclosure, the calculation order and target algorithm of multiple kernel functions in the initial model are first determined, and then the multiple kernel functions are classified. Then, based on the classification results and the calculation order, the target subgraph can be determined, and finally the initial model is optimized using the target algorithm and the target subgraph, thereby achieving the purpose of being able to use the optimization algorithm for some kernel functions in the model, solving the technical problem in the related technology that when some kernel functions do not match the optimization algorithm, the optimization algorithm will still optimize the unmatched kernel functions, and achieving the effect of being able to use the optimization algorithm for the kernel functions in the model that match the optimization algorithm and removing the kernel functions that do not match the optimization algorithm, thereby solving the technical problem of low optimization efficiency for model optimization

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a model processing method according to an embodiment of the present disclosure;

[0013] Figure 2 is a flow chart of a model processing method according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram of an optional subgraph optimization according to an embodiment of the present disclosure;

[0015] Figure 4 It is a structural block diagram of a processing device of a model according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0016] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0017] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0018] Existing technical solutions typically optimize the entire model or a single kernel function. This approach often has some drawbacks. For example, using a static graph optimization algorithm (cuda_graph) with long sequences of long short-term memory (LSTM) can impose a significant burden on graphics memory. The removal / addition of padding operations between consecutive attention mechanism calculations can burden data inflow and outflow (IO), slowing down computations.

[0019] Overall optimization plan: Due to various reasons, some optimization plans may not be applicable to certain operators (ops), resulting in the inability to optimize the entire model. For example, in the whole sentence mode, the LSTM for long speech recognition generates an overly large cuda_graph due to loop unrolling, which consumes GPU resources and affects the final model launch.

[0020] Kernel optimization: Some optimization schemes are not suitable for optimizing a specific operation individually. For example, in densely packed algorithms, adding padding removal / restoration operations before and after the kernel operation to be optimized will result in a large number of read and write operations, increasing the model's forward inference time.

[0021] According to an embodiment of the present disclosure, a method for processing a model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0022] The method embodiments provided in the embodiments of the present disclosure can be executed in a mobile terminal, a computer terminal or a similar electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a model processing method is shown.

[0023] like Figure 1 As shown, the computer terminal 100 includes a computing unit 101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 102 or a computer program loaded from a storage unit 108 into a random access memory (RAM) 103. Various programs and data required for the operation of the computer terminal 100 can also be stored in the RAM 103. The computing unit 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.

[0024] Multiple components in the computer terminal 100 are connected to the I / O interface 105, including an input unit 106, such as a keyboard, a mouse, etc.; an output unit 107, such as various types of displays, speakers, etc.; a storage unit 108, such as a magnetic disk, an optical disk, etc.; and a communication unit 109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 109 allows the computer terminal 100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0025] The computing unit 101 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 101 performs the processing method of the model described herein. For example, in some embodiments, the processing method of the model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer terminal 100 via ROM 102 and / or communication unit 109. When the computer program is loaded into RAM 103 and executed by the computing unit 101, one or more steps of the processing method of the model described herein can be performed. Alternatively, in other embodiments, the computing unit 101 can be configured to execute the processing method of the model by any other appropriate means (e.g., by means of firmware).

[0026] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0027] It should be noted that, in some optional embodiments, the above Figure 1 The electronic device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the electronic device described above.

[0028] Under the above operating environment, the present disclosure provides the following Figure 2 The processing method of the model shown in Figure 1The computer terminal or similar electronic device shown is used for execution. Figure 2 Flowchart of a model processing method provided according to an embodiment of the present disclosure. Figure 2 As shown, the method may include the following steps:

[0029] Step S20: Obtain the calculation order and target algorithm of multiple kernel functions in the initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model.

[0030] The above-mentioned initial model can be any one or more models that need to be optimized; the kernel function can be any one or more kernel functions that can be calculated, for example, it can be a linear kernel function, a Gaussian kernel function, an exponential kernel function, etc. The specific kernel function can be set according to user needs and is not specifically limited in this embodiment. Among them, the kernel function can make the linearly inseparable pattern in the low-dimensional space be nonlinearly mapped to the high-dimensional feature space, thereby making the linearly inseparable pattern in the low-dimensional space linearly separable.

[0031] The above calculation order can be one or more orders used to represent the input-output relationship between kernel functions. For example, if the output of kernel function K1 is the input of kernel function K2, the calculation order can be to calculate kernel function K1 first and then calculate kernel function K2, but it is not limited to this.

[0032] The above-mentioned target algorithm can be any one or more algorithms that can optimize the initial model, for example, it can be a half-precision calculation algorithm, a mixed-precision calculation algorithm, a tensorcore optimization algorithm, a parallel instruction set optimization algorithm, etc. The specific optimization algorithm can be set according to user needs. In this embodiment, the dense-packed algorithm and the cuda_graph algorithm are taken as examples. Among them, the dense-packed algorithm can merge requests of different lengths together for inference, which can avoid the calculation of fillers and reduce the amount of calculation. After creating the cuda_graph algorithm, the central processing unit (CPU) can be avoided from sending each step of the calculation command all the time, and the created cuda_graph algorithm can be directly called.

[0033] The above-mentioned calculation operations can be any one or more operations that can calculate the initial model through the kernel function. For example, it can be selecting a static graph instance when using the cuda_graph algorithm, removing padding from the input of each kernel function, mixed precision calculation optimization, half precision calculation optimization, etc., but is not limited to these.

[0034] In an optional embodiment, the calculation order and target algorithm of multiple kernel functions in the initial model can be obtained first. For example, the input-output relationship between multiple kernel functions K1, K2, and K3 and the dense-packed algorithm and cuda_graph algorithm corresponding to the multiple kernel functions can be obtained first. It should be noted that kernel functions are not limited to K1-K3, and the target algorithm is not limited to the dense-packed algorithm and cuda_graph algorithm. In this embodiment, K1-K3, the dense-packed algorithm, and the cuda_graph algorithm are used as examples for explanation. In this step, by obtaining the calculation order and target algorithm of multiple kernel functions in the initial model, a basis for subsequent classification of the kernel functions can be provided.

[0035] Step S21 , classifying the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm.

[0036] The above classification result may be that the kernel function supports the target algorithm or that the kernel function does not support the target algorithm, but is not limited thereto. Support may indicate whether the kernel function can be optimized by the target algorithm.

[0037] In an optional embodiment, the kernel function can be classified by adding an auxiliary flag flag_support to the kernel function to determine whether multiple kernel functions in the initial model support the target algorithm, wherein flag_support can be set by oneself; optionally, flag_support can be set by the user; optionally. flag_support can be a mark set by the user to indicate whether a kernel function supports the target algorithm, wherein flag_support can be false to indicate that the kernel function does not support the target algorithm, and flag_support can be true to indicate that the kernel function supports the target algorithm, but is not limited to this. In this step, by classifying multiple kernel functions, the kernel function that supports the target algorithm can be obtained, which can provide a basis for subsequent subgraph optimization.

[0038] Step S22: determining a target subgraph based on the classification result and the calculation order, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function supporting the target algorithm.

[0039] The above-mentioned target subgraph can be a subgraph that supports the optimization algorithm in any one or more initial models, wherein a graph G can be initialized, and there are N points in the graph G, each point corresponding to each kernel function in the initial model, and each point contains a flag corresponding to flag_support of each kernel function in the initial model. Then the graph G is the parent graph corresponding to the initial model, and the graph obtained based on the classification results and input-output relationship (i.e., calculation order) of multiple kernel functions is the target subgraph of the initial model, wherein the target subgraph is included in the parent graph. In the form of a subgraph, the code development work of developing the kernel separately is avoided, making the subsequent access to new kernels more convenient.

[0040] The above calculation order can be determined by adding directed edges to each point in the mother graph, so as to determine the input-output relationship of each point (i.e., each kernel function in the initial model). For example, a directed edge pointing to point n1 and a directed edge starting from it can be added to point n1, and the output of point n0 that sends a directed edge pointing to point n1 can be indicated as the input of point n1, so as to determine the calculation order of each point (i.e., each kernel function in the initial model). Among them, the edge between the fixed point Ni and Nj has a direction, and this edge is called a directed edge.

[0041] The above operation process can be a process of inputting the data of the kernel function into the next kernel function through the directed edge of the point in the target subgraph (i.e., the kernel function in the initial model), and the data string composed of the input and output of multiple kernel functions.

[0042] In an optional embodiment, first, a graph G can be initialized. There are N points in the graph G, each point corresponds to each kernel function in the initial model, each point contains a flag bit, and each flag bit corresponds to the flag_support of each kernel function in the initial model. Then the graph G is the mother graph corresponding to the initial model. Secondly, a directed edge can be added to each point to determine the input-output relationship of each point. Finally, based on the flag bit of each point (that is, the flag_support of each kernel function in the initial model) and the input-output relationship, the subgraph in the initial model that supports the optimization algorithm can be determined.

[0043] In this step, by establishing the corresponding mother graph for the initial model and determining the target subgraph through the classification results and calculation order, a basis can be provided for the subsequent optimization of the target subgraph.

[0044] Step S23: Optimize the initial model using the target algorithm and the target subgraph to obtain the target model.

[0045] In an optional embodiment, first, the calculation operation corresponding to the target algorithm can be determined, and secondly, the initial model can be optimized based on the calculation operation and the target subgraph, that is, the optimized subgraph (i.e., the target model) can be obtained; for example, first, the target algorithm can be selected as the close-packed algorithm, and secondly, the calculation operation corresponding to the close-packed algorithm can be determined, and finally, the points in the initial model that do not support the optimization algorithm (i.e., each kernel function in the initial model) can be removed through the calculation operation and the target subgraph, and the optimized subgraph (i.e., the target model) can be obtained. It should be noted that the target algorithm in this embodiment is not limited to the close-packed algorithm. In this embodiment, the close-packed algorithm is taken as an example for explanation.

[0046] In another optional embodiment, first, the calculation operation corresponding to the target algorithm can be determined, and then the initial model can be optimized based on the calculation operation and the target subgraph, that is, the optimized subgraph (i.e., the target model) can be obtained; for example, the target algorithm can first be selected as the cuda_graph algorithm, and then the calculation operation corresponding to the cuda_graph algorithm can be determined. Finally, the points in the initial model that do not support the optimization algorithm (i.e., each kernel function in the initial model) can be removed through the calculation operation and the target subgraph, and the optimized subgraph (i.e., the target model) can be obtained. It should be noted that the target algorithm in this embodiment is not limited to the cuda_graph algorithm. In this embodiment, the cuda_graph algorithm is used as an example for explanation. In this step, by optimizing the initial model using the target algorithm and the target subgraph, the technical problem of low algorithm optimization rate caused by problems with a kernel function of the model is solved.

[0047] Optionally, the target model is a speech processing model, and the method further includes: processing the target speech using the speech processing model to obtain a speech processing result.

[0048] The target speech mentioned above may be any one or more speech that needs to be processed.

[0049] In an optional embodiment, the target model may be a speech processing model, which may implement functions such as speech recognition, speech detection, speech classification, and speech extraction.

[0050] In the embodiments of the present disclosure, speech extraction is taken as an example for illustration. When it is necessary to extract the target speech in a speech segment, the target speech can be extracted by the speech processing model in the embodiments of the present disclosure, that is, the required target speech information, that is, the speech processing result, can be obtained.

[0051] In another optional embodiment, speech recognition is used as an example for illustration, wherein, when speech recognition is required for a target speech in a speech segment, the target speech can be recognized by the speech processing model in the embodiment of the present disclosure, that is, the required recognition result of the target speech can be obtained, that is, the speech processing result, wherein the recognition result can be used to represent the specific content of the target speech.

[0052] According to the above steps S20 to S23 of the present invention, first, the calculation order and target algorithm of multiple kernel functions in the initial model are determined, then the multiple kernel functions are classified, and then the target subgraph can be determined based on the classification results and the calculation order. Finally, the initial model is optimized based on the target algorithm and the target subgraph, thereby achieving the purpose of being able to use the optimization algorithm for some kernel functions in the model, solving the technical problem in the related technology that when some kernel functions do not match the optimization algorithm, the optimization algorithm will still optimize the unmatched kernel functions, and achieving the effect of being able to use the optimization algorithm for the kernel functions in the model that match the optimization algorithm and removing the kernel functions that do not match the optimization algorithm, thereby solving the technical problem of low optimization efficiency for optimizing the model.

[0053] The above method of this embodiment is further introduced below.

[0054] Optionally, the target subgraph is optimized using a target algorithm to obtain a target model, including: determining a target operation corresponding to the target algorithm, wherein the target operation is used to represent the corresponding operation when the target algorithm optimizes the target subgraph; and optimizing the initial model based on the target operation and the target subgraph to obtain the target model.

[0055] The above-mentioned target operation can be an operation for optimizing the target model through the target algorithm and the target subgraph. For example, when the target algorithm is a dense packing algorithm, the target operation can be to remove the padding from the input of the initial model through the dense packing algorithm and the target subgraph to achieve dense packing; when the target algorithm is the cuda_graph algorithm, the target operation can be to select the static graph instance cuda_graph_instance through the cuda_graph algorithm and the target subgraph when calling cuda_graph, thereby reducing the calculation commands sent by the CPU. It should be noted that the target algorithm in this embodiment is not limited to the dense packing algorithm and the cuda_graph algorithm. In this embodiment, the dense packing algorithm and the cuda_graph algorithm are taken as examples for illustration.

[0056] In an optional embodiment, a target operation corresponding to a target algorithm may be first determined, and then the initial model may be optimized based on the target operation and the target subgraph. For example, padding may be removed from the input of the initial model to achieve dense packing using the target operation and the target subgraph corresponding to the dense packing algorithm, thereby obtaining an optimized subgraph (i.e., the target model). However, the present invention is not limited thereto. In this step, the initial model is optimized using the optimization algorithm and the target subgraph to obtain an optimized subgraph.

[0057] Optionally, when the target algorithm is a static graph optimization algorithm, the initial model is optimized based on the target operation and the target subgraph to obtain the target model, including: encapsulating the target subgraph into a target static graph; calling the target static graph using the target interface to obtain a calling result; and optimizing the initial model based on the calling result and the target subgraph to obtain the target model.

[0058] The target static graph may be a static graph that needs to be optimized; the target interface may be a cuda_graph interface, but is not limited thereto; and the call result may be a call success or a call failure, but is not limited thereto.

[0059] In an optional embodiment, when the target algorithm is the cuda_graph algorithm (i.e., a static graph optimization algorithm), the target subgraph can first be encapsulated as a target static graph, and then the target static graph can be called through the cuda_graph interface (i.e., the target interface). When the call is successful, the initial model can be optimized by sending the static graph instance cuda_graph_instance and the target subgraph, and the optimized subgraph (i.e., the target model) can be obtained. The target subgraph can be encapsulated into a target static graph through code encapsulation, but is not limited to this. Encapsulation can make the call-related code more concise. In this step, the initial model is optimized through the target static graph, the target interface, and the target subgraph, which can reduce the calculation instructions sent by the CPU.

[0060] Optionally, when the target algorithm is a dense-packed optimization algorithm, the initial model is optimized based on the target operation and the target subgraph to obtain the target model, including: obtaining initial input characters, wherein the initial input characters are input characters to be input into the target subgraph; deleting the target symbols in the initial input characters to obtain target input characters; inputting the target input characters into the target subgraph to output the initial output characters; adding the target symbols to the initial output characters to obtain the target output characters; optimizing the initial model based on the target input characters, the target output characters and the target subgraph to obtain the target model.

[0061] The above-mentioned initial input characters can be input characters that need to be input into the target subgraph, for example, they can be 1, 2, 3, m, n, k, x, y, j, etc., but not limited to this; the target symbol can be a symbol filled to make the request lengths of different lengths equal, for example, it can be "~", but not limited to this; the target input character can be the initial input character after deleting the target symbol to reduce the calculation scale of the target subgraph; the initial output character can be the output character output by the target subgraph, for example, it can be 1, 2, 3, m, n, k, x, y, j, etc., but not limited to this, because the part of the initial model other than the target subgraph does not support the dense packing algorithm, it is necessary to add the target character to the initial output character so that the initial model can use the optimization algorithm normally; the target output character can be the character after adding the target symbol to the initial output character.

[0062] In an optional embodiment, when the target algorithm is a dense packing optimization algorithm, the initial input characters of the target subgraph can be obtained first, and then the target characters in the initial input characters are removed by the padding operation in the dense packing algorithm to obtain the target input characters, so as to reduce the calculation scale of the target subgraph, and then the target input characters are input to the target subgraph, so that the target subgraph can output the initial output characters. Because the parts other than the target subgraph in the initial model do not support the dense packing algorithm, it is necessary to add the target characters to the initial output characters to obtain the target output characters so that the initial model can use the optimization algorithm normally. Finally, the points in the initial model that do not support the optimization algorithm can be removed by the target input characters, the target output characters, and the target subgraph, and the optimized subgraph (i.e., the target model) can be obtained. In this step, the calculation scale of the target subgraph can be reduced.

[0063] Optionally, based on the classification results and the calculation order, a target subgraph of the target kernel function is determined, including: based on the classification results and the calculation order, a target directed graph is determined, wherein the target directed graph includes multiple nodes and at least one directed edge, the multiple nodes correspond to multiple kernel functions, and the at least one directed edge is used to represent the calculation order of the multiple kernel functions; based on the first node, an initial connected domain of the target directed graph is processed to obtain a target connected domain, wherein the kernel function corresponding to the first node does not support the target algorithm, and the nodes contained in the target connected domain correspond to the target kernel function; based on the target connected domain, a target subgraph is determined.

[0064] The above-mentioned target directed graph can be a directed graph that needs to be optimized, which can be represented as a graph G, wherein the graph G contains multiple nodes and at least one directed edge, each node corresponds to each kernel function in the initial model, and each node contains a flag bit, corresponding to the flag_support of the kernel function in the initial model, so that the flag_support of the nodes in the graph G that support the target algorithm can be true, and the flag_support of the nodes that do not support the target algorithm can be false; the directed edges can represent the calculation order of multiple kernel functions. For example, the directed edge of node n1 points to node n2, which means that the output of node n1 is the input of node n2, then the calculation order can be to calculate node n1 first, and then calculate node n2. It should be noted that the nodes in the directed graph are not limited to n1 and n2. In this embodiment, n1 and n2 are used as examples for explanation.

[0065] The above-mentioned initial connected domain can be a connected domain that has not been optimized, in which all nodes and directed edges are connected to each other to form a set, called a connected domain; the target connected domain can be a connected domain after removing nodes that do not support the optimization algorithm, that is, an optimized connected domain, in which the target connected domain contains nodes with flag_support set to true (that is, the target kernel function).

[0066] The first node mentioned above may be a node that does not support the target algorithm, that is, may be a node where flag_support is false.

[0067] In an optional embodiment, first, based on the classification results and the calculation order, the graph G (i.e., the target directed graph) can be determined, wherein the graph G contains multiple nodes and at least one directed edge, each node corresponds to each kernel function in the initial model, and each node contains a flag bit, corresponding to the flag_support of the kernel function in the initial model, so that the flag_support of the nodes in the graph G that support the target algorithm can be true, and the flag_support of the nodes that do not support the target algorithm can be false, and each directed edge can represent the calculation order between each node; secondly, the nodes with flag_support being false in the initial connected domain of the target directed graph are deleted to obtain the target connected domain, wherein the target connected domain contains nodes with flag_support being true (i.e., the target kernel function) and directed edges connected by the nodes with flag_support being true; finally, the target subgraph can be determined based on the target connected domain, wherein the target subgraph contains the target connected domain. In this step, the subgraph supporting the optimization algorithm (ie, the target subgraph) can be determined through the target directed graph, the target connected domain, and the first node, which can provide a basis for subsequent optimization of the target subgraph.

[0068] Optionally, the initial connected domain of the target directed graph is processed based on the first node to obtain the target connected domain, including: in response to the initial connected domain containing the first node and the existence of a directed edge of the first node, determining the first directed edge corresponding to the first node; deleting the first directed edge in the initial connected domain to obtain a first connected domain, wherein the first connected domain contains a second node and the kernel function corresponding to the second node supports the target algorithm; processing the first connected domain based on the first node to obtain the target connected domain.

[0069] The above-mentioned first directed edge can be a directed edge corresponding to a node that does not support the target algorithm (i.e., the first node). In this embodiment, the first directed edge includes at least: a directed edge pointing to the first node, and a directed edge issued by the first node; the first connected domain can be a connected domain that includes a kernel function that supports the target algorithm (i.e., the second node).

[0070] The second node mentioned above may be a kernel function in the initial model that supports the target algorithm.

[0071] In an optional embodiment, first, in response to the initial connected domain containing a node that does not support the target algorithm (i.e., the first node), and the first node having a connected directed edge, the first directed edge corresponding to the first node can be deleted. For example, when the first directed edge corresponding to the first node is a directed edge pointing to the first node, the first directed edge is deleted, and a first connected domain can be obtained, wherein the first connected domain contains nodes that support the target algorithm. Secondly, the first node in the first connected domain is deleted, and the target connected domain can be obtained.

[0072] In another optional embodiment, first, in response to the initial connected domain containing a first node and the first node having a connected directed edge, the first directed edge corresponding to the first node can be deleted. For example, when the first directed edge corresponding to the first node is a directed edge issued by the first node, the first directed edge is deleted to obtain a first connected domain, wherein the first connected domain contains nodes that support the target algorithm. Secondly, the first node in the first connected domain is deleted to obtain the target connected domain.

[0073] In another optional embodiment, first, in response to the initial connected domain containing a first node and the first node having a connected directed edge, the first directed edge corresponding to the first node can be deleted. For example, when the first directed edge corresponding to the first node is a directed edge issued by the first node and a directed edge pointing to the first node, the first directed edge is deleted to obtain a first connected domain, wherein the first connected domain contains nodes that support the target algorithm. Secondly, the first node in the first connected domain is deleted to obtain the target connected domain. In this step, the target connected domain that supports the target algorithm can be obtained by deleting the directed edge of the first node in the initial connected domain of the target directed graph.

[0074] Optionally, processing the initial connected domain of the target directed graph based on the first node to obtain the target connected domain includes: in response to the initial connected domain not including the first node, directly determining the initial connected domain as the target connected domain.

[0075] In an optional embodiment, when the initial connected domain of the target directed graph does not include a node that does not support the target algorithm (i.e., the first node), the initial connected domain can be directly determined as the target connected domain. In this step, by determining that the initial connected domain does not include the first node, the target connected domain that includes the target algorithm can be directly obtained.

[0076] Optionally, the first connected domain is processed based on the first node to obtain a target connected domain, including: in response to the existence of the target node in the first connected domain, updating the target node in the first connected domain to the first node to obtain a second connected domain, wherein the target node corresponds to the node pointed to by the first directed edge in the initial connected domain; determining the first directed edge corresponding to the first node; and deleting the first directed edge in the second connected domain to obtain the target connected domain.

[0077] The target node mentioned above may be the node pointed to by the first directed edge of the node (ie, the first node) that does not support the target algorithm in the first connected domain.

[0078] The second connected domain mentioned above may be a connected domain obtained by updating the node pointed to by the first directed edge to the first node.

[0079] In an optional embodiment, in response to the presence of a node (i.e., a target node) pointed to by the first directed edge in the first connected domain, the target node can be updated to the first node, thereby obtaining a second connected domain. Next, the first directed edge corresponding to the first node is determined and deleted to obtain the target connected domain. In this step, the target connected domain containing the target algorithm is obtained by deleting the directed edge of the updated first node.

[0080] Optionally, processing the first connected domain based on the first node to obtain the target connected domain includes: in response to the first connected domain not including the target node, directly determining the first connected domain as the target connected domain.

[0081] In an optional embodiment, it may be first determined whether the first connected domain contains the target node. In response to the first connected domain not containing the target node, the first connected domain may be directly determined as the target connected domain. In this step, by determining that the first connected domain does not contain the first node, the target connected domain containing the target algorithm may be directly obtained.

[0082] Optionally, multiple kernel functions are classified to obtain classification results, including: obtaining a target identifier corresponding to each kernel function, wherein the target identifier is used to indicate whether each kernel function supports a target algorithm; and classifying multiple kernel functions based on the target identifier to obtain classification results.

[0083] The above-mentioned target identifier can be any one or more identifiers that can indicate whether each kernel function supports the target algorithm. In this embodiment, the auxiliary flag bit flag_support is taken as an example. For example, if the kernel function K1 supports the target algorithm, the auxiliary flag bit corresponding to K1 can be true. If the kernel function K2 does not support the target algorithm, the auxiliary flag bit corresponding to K2 can be false.

[0084] In an optional embodiment, first, an auxiliary identification bit (i.e., a target identifier) ​​can be added to each kernel function in the initial model. Then, the target identifier corresponding to each kernel function is obtained, and multiple kernel functions are classified based on the target identifier to obtain a classification result. For example, the kernel function can be classified based on flag_support to obtain a classification result of whether the kernel function supports the target algorithm. In this step, the kernel function can be classified and the classification result can be obtained through the auxiliary identification bit of the kernel function.

[0085] Optionally, the target model is an image processing model, and the method further includes: processing the target image using the image processing model to obtain an image processing result.

[0086] The target image mentioned above may be any one or more images that require image processing.

[0087] In an optional embodiment, the target model may also be an image processing model, which can realize functions such as image recognition, image detection, image classification, and image extraction.

[0088] In the embodiments of the present disclosure, image recognition is taken as an example for illustration. For example, when a target image needs to be identified, the target image can be identified through the image processing model of the embodiments of the present disclosure, and a recognition result of the target image, that is, an image processing result, can be obtained. The recognition result can be used to represent the specific content in the target image.

[0089] In another optional embodiment, image classification is used as an example. For example, when a target image needs to be classified, the target image can be classified using the image processing model in the embodiment of the present disclosure to obtain the category of the target image, that is, the image processing result.

[0090] Optionally, the target model is a natural language processing model, and the method further includes: processing the target text using the natural language processing model to obtain a text processing result.

[0091] The above-mentioned natural language processing model can be any one or more models that need to be processed.

[0092] In an optional embodiment, the target model may also be a natural language processing model, which can implement functions such as text recognition, text detection, text classification, and text extraction.

[0093] In the embodiments of the present disclosure, text detection is taken as an example for illustration. For example, when it is necessary to detect the target text, the target text can be detected by the natural language processing model of the embodiments of the present disclosure, and the detection result of the target text, that is, the text processing result, can be obtained.

[0094] In another optional embodiment, text extraction is taken as an example. For example, when it is necessary to extract the target field of the target text, the target field can be extracted through the natural language processing model in the embodiment of the present disclosure, and the target field information, that is, the text processing result, can be obtained.

[0095] The present disclosure proposes a local optimization algorithm, the main steps of which are as follows:

[0096] Step S1, specifying the optimization algorithm to be used, and setting the optimization algorithm to be Algorithm, where Algorithm can represent a dense packing algorithm or a cuda_graph algorithm;

[0097] In step S2, an auxiliary flag flag, flag_support, is added to the kernel. Since the specific kernel execution is known, adding this flag can clarify whether this type of kernel is suitable for the optimization algorithm in step S1. For example, in the whole sentence mode, very long sentences are not suitable for the cuda_graph algorithm. Calculate softmax(q*k^T)*v. Because the Basic Linear Algebra Subprograms (ALBS) programming language (C language) interface is used, the dense packing algorithm is not applicable. Softmax(q*k^T)*v represents the calculation formula of the attention mechanism.

[0098] Step S3: Based on step S2, let the model have N kernels and M directed edges. The directed edges indicate that the output of the starting kernel is the input of the target kernel.

[0099] Step S4: Initialize a graph G with N points. Each point in the graph corresponds to a kernel in the model. Each point contains a flag bit corresponding to the kernel's flag_support. The graph has M directed edges (E), corresponding to the M directed edges in the model structure. Set flag_support of points n1, n2, ..., nk to false, and set flag_support of the remaining points to true.

[0100] Step S5, set a set S, where S is initialized to a set of k points with flag_support set to false;

[0101] Step S6: If S is empty, proceed to step S9; if S is not empty, take a node from S, let it be node n_i, and delete all directed edges pointing to n_i from G;

[0102] Step S7, traverse all directed edges E(n_i) starting from node n_i, let one of the directed edges be e_y, delete e_y from G, set flag_support of n_j pointed to by e_y to false, and put n_j into S;

[0103] Step S8, jump back to step 6;

[0104] Step S9: Calculate the connected domains of nodes in G where flag_support is true at this time. Let G have K connected domains. Let one of these connected domains be K_x. K_x corresponds to the computation of a portion of the kernel in the model structure. Let the computation be represented as F(X). This portion of the computation can be optimized using an algorithm. F(X) is called a subgraph. This subgraph is optimized.

[0105] Optimizing some subgraphs includes the following steps:

[0106] Step S1, read the model structure and determine the data input and output relationship between each kernel;

[0107] Step S2: Calculate the set of subgraphs according to the above-mentioned local optimization algorithm, and let the subgraph set be K;

[0108] In step S3, the subgraph portion is optimized, while the non-subgraph portion is not optimized. The optimization operation may be the cuda_graph algorithm or the dense packing algorithm. It should be noted that the optimization operation in this embodiment may be any operation capable of optimizing a subgraph. In this embodiment, the cuda_graph algorithm and the dense packing algorithm are used as examples for illustration.

[0109] a) The Cuda_graph algorithm, when not optimized, sends kernel instructions sequentially. After optimization, for each subgraph in K, it encapsulates it into a cuda_graph, and selects the cuda_graph execution instruction cuda_graph_instance when calling cuda_graph, while the non-subgraph part sends kernel execution instructions.

[0110] b) Dense packing algorithm: remove padding from the input of each subgraph in K to achieve dense packing. Since the dense packing algorithm cannot be used for non-subgraph parts, padding is added to the output of each subgraph so that the initial model can use the optimization algorithm normally.

[0111] Figure 3 is a schematic diagram of an optional subgraph optimization according to an embodiment of the present disclosure, such as Figure 3 As shown, S1 represents the structure of the initialization graph, the number represents the sequence number of the node, and false / true represents whether the optimization algorithm is supported. At this time, set S = (2), S2 represents removing node 2 from set S, and the main step is to delete the directed edge pointing to node 2 and the directed edge emitted from node 2. At this time, set S = (4), S3 represents removing node 4 from set S, and the main step is to delete the directed edge pointing to node 4 and the directed edge emitted from node 4, and put the node pointed to by the directed edge emitted from node 4 into set S. At this time, set S is empty, and the subgraph division step is completed.

[0112] The beneficial effects of the present disclosure are:

[0113] 1. Reduce development effort: By using subgraphs, the code development work of developing a separate kernel is avoided, making it easier to integrate new kernels later.

[0114] 2. Support multiple optimization algorithms: When the entire model is not suitable for a certain optimization algorithm, the model can be locally optimized to improve the flexibility of forward calculation and reduce calculation time.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.

[0116] The present disclosure also provides a model processing device for implementing the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0117] Figure 4 is a structural block diagram of a processing device of a model according to an embodiment of the present disclosure, such as Figure 4 As shown, a model processing device 400 includes: an acquisition module 401, used to obtain the calculation order and target algorithm of multiple kernel functions in the initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model; a classification module 402, used to classify the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm; a determination module 403, used to determine the target subgraph based on the classification results and the calculation order, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function that supports the target algorithm; an optimization module 404, used to optimize the initial model using the target algorithm and the target subgraph to obtain the target model.

[0118] Optionally, the optimization module includes: a first determination unit, used to determine a target operation corresponding to a target algorithm, wherein the target operation is used to represent the corresponding operation when the target algorithm optimizes a target subgraph; an optimization unit, used to optimize the initial model based on the target operation and the target subgraph to obtain a target model.

[0119] Optionally, when the target algorithm is a static graph optimization algorithm, the optimization unit includes: an encapsulation sub-unit, used to encapsulate the target sub-graph into a target static graph; a calling sub-unit, used to call the target static graph using the target interface to obtain a calling result; and a first optimization sub-unit, used to optimize the initial model based on the calling result and the target sub-graph to obtain a target model.

[0120] Optionally, when the target algorithm is a dense optimization algorithm, the optimization unit also includes: an acquisition subunit, used to acquire initial input characters, wherein the initial input characters are input characters to be input into the target subgraph; a deletion subunit, used to delete the target symbol in the initial input characters to obtain the target input characters; an input subunit, used to input the target input characters into the target subgraph to output the initial output characters; an addition subunit, used to add the target symbol to the initial output characters to obtain the target output characters; and a second optimization subunit, used to optimize the initial model based on the target input characters, the target output characters and the target subgraph to obtain the target model.

[0121] Optionally, the determination module includes: a first determination unit, used to determine the target directed graph based on the classification results and the calculation order, wherein the target directed graph includes multiple nodes and at least one directed edge, the multiple nodes correspond to multiple kernel functions, and the at least one directed edge is used to represent the calculation order of the multiple kernel functions; a processing unit, used to process the initial connected domain of the target directed graph based on the first node to obtain the target connected domain, wherein the kernel function corresponding to the first node does not support the target algorithm, and the nodes contained in the target connected domain correspond to the target kernel function; a second determination unit, used to determine the target subgraph based on the target connected domain.

[0122] Optionally, the processing unit includes: a first determination subunit, used to determine the first directed edge corresponding to the first node in response to the initial connected domain containing a first node and the existence of a directed edge in the first node; a deletion subunit, used to delete the first directed edge in the initial connected domain to obtain a first connected domain, wherein the first connected domain contains a second node and the kernel function corresponding to the second node supports the target algorithm; a processing subunit, used to process the first connected domain based on the first node to obtain a target connected domain.

[0123] Optionally, the processing unit further includes: a second determining subunit, configured to directly determine the initial connected domain as the target connected domain in response to the initial connected domain not including the first node.

[0124] Optionally, the processing sub-unit is further used to, in response to the existence of a target node in the first connected domain, update the target node in the first connected domain to the first node to obtain a second connected domain, wherein the target node corresponds to the node pointed to by the first directed edge in the initial connected domain; determine the first directed edge corresponding to the first node; delete the first directed edge in the second connected domain to obtain the target connected domain.

[0125] Optionally, the processing subunit is further configured to, in response to the first connected domain not including the target node, directly determine the first connected domain as the target connected domain.

[0126] Optionally, the classification module includes: an acquisition unit for acquiring a target identifier corresponding to each kernel function, wherein the target identifier is used to indicate whether each kernel function supports a target algorithm; and a classification unit for classifying multiple kernel functions based on the target identifier to obtain a classification result.

[0127] Optionally, the target model is a speech processing model, and the device further includes: a first processing model, configured to process the target speech using the speech processing model to obtain a speech processing result.

[0128] Optionally, the target model is an image processing model, and the device further includes: a second processing model, configured to process the target image using the image processing model to obtain an image processing result.

[0129] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0130] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including a memory and at least one processor, wherein the memory stores computer instructions, and the processor is configured to execute the computer instructions to perform the steps in any of the above method embodiments.

[0131] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0132] Optionally, in the present disclosure, the processor may be configured to execute the following steps through a computer program:

[0133] S1, obtaining the calculation order and target algorithm of multiple kernel functions in the initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model;

[0134] S2, classifying the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm;

[0135] S3, based on the classification results and the calculation order, determining the target subgraph, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function that supports the target algorithm;

[0136] S4, optimize the initial model using the target algorithm and target subgraph to obtain the target model.

[0137] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0138] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are configured to execute the steps of any of the above method embodiments during runtime.

[0139] Optionally, in this embodiment, the non-volatile storage medium may be configured to store a computer program for executing the following steps:

[0140] S1 obtains the calculation order and target algorithm of multiple kernel functions in the initial model, wherein the multiple kernel functions are used to represent multiple calculation operations of the initial model;

[0141] S2, classifying the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm;

[0142] S3, based on the classification results and the calculation order, determining the target subgraph, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function that supports the target algorithm;

[0143] S4, optimize the initial model using the target algorithm and target subgraph to obtain the target model.

[0144] Alternatively, in this embodiment, the non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any suitable combination of the above. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0145] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product. The program code for implementing the method embodiment of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0147] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0148] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0149] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0151] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.

Claims

1. A model processing method, comprising: Obtaining a calculation order and a target algorithm for a plurality of kernel functions in an initial model, wherein the plurality of kernel functions are used to represent a plurality of calculation operations of the initial model; Classifying the multiple kernel functions to obtain classification results, wherein the classification results are used to indicate whether the multiple kernel functions support the target algorithm; Determining a target subgraph based on the classification result and the calculation order, wherein the target subgraph is used to represent the operation process of the target kernel function, and the target kernel function is used to represent the kernel function supporting the target algorithm; Optimizing the initial model using the target algorithm and the target subgraph to obtain a target model, wherein the target model is a speech processing model or an image processing model; The method further includes: processing the target speech using the speech processing model to obtain a speech processing result; or processing the target image using the image processing model to obtain an image processing result; The determining of the target subgraph based on the classification result and the calculation order includes: determining a target directed graph based on the classification result and the calculation order, wherein the target directed graph includes a plurality of nodes and at least one directed edge, the plurality of nodes corresponding to the plurality of kernel functions, and the at least one directed edge is used to represent the calculation order of the plurality of kernel functions; processing an initial connected domain of the target directed graph based on a first node to obtain a target connected domain, wherein the kernel function corresponding to the first node does not support the target algorithm, and the nodes included in the target connected domain correspond to the target kernel function; and determining the target subgraph based on the target connected domain; The method processes an initial connected domain of the target directed graph based on the first node to obtain a target connected domain, including: in response to the initial connected domain containing the first node and the first node having a directed edge, determining a first directed edge corresponding to the first node; deleting the first directed edge in the initial connected domain to obtain a first connected domain, wherein the first connected domain contains a second node and the kernel function corresponding to the second node supports the target algorithm; and processing the first connected domain based on the first node to obtain the target connected domain.

2. The method according to claim 1, wherein the initial model is optimized using the target algorithm and the target subgraph to obtain a target model, comprising: Determining a target operation corresponding to the target algorithm, wherein the target operation is used to represent an operation corresponding to when the target algorithm optimizes the target subgraph; The initial model is optimized based on the target operation and the target subgraph to obtain the target model.

3. The method according to claim 2, wherein when the target algorithm is a static graph optimization algorithm, optimizing the initial model based on the target operation and the target subgraph to obtain the target model comprises: Encapsulating the target subgraph into a target static graph; Using the target interface to call the target static graph, and obtain a call result; The initial model is optimized based on the call result and the target subgraph to obtain the target model.

4. The method according to claim 2, when the target algorithm is a dense packing optimization algorithm, optimizing the initial model based on the target operation and the target subgraph to obtain the target model, comprising: Acquire an initial input character, wherein the initial input character is an input character to be input into the target subgraph; Deleting the target symbol from the initial input character to obtain the target input character; Inputting the target input character into the target subgraph, and outputting an initial output character; Adding the target symbol to the initial output character to obtain a target output character; The initial model is optimized based on the target input character, the target output character and the target subgraph to obtain the target model.

5. The method according to claim 1, wherein the initial connected domain of the target directed graph is processed based on the first node to obtain the target connected domain, comprising: In response to the initial connected domain not including the first node, the initial connected domain is directly determined as the target connected domain.

6. The method according to claim 1, wherein processing the first connected domain based on the first node to obtain the target connected domain comprises: In response to the existence of a target node in the first connected domain, updating the target node in the first connected domain to the first node to obtain a second connected domain, wherein the target node corresponds to the node pointed to by the first directed edge in the initial connected domain; Determine the first directed edge corresponding to the first node; The first directed edge in the second connected domain is deleted to obtain the target connected domain.

7. The method according to claim 6, wherein processing the first connected domain based on the first node to obtain the target connected domain comprises: In response to the first connected domain not including the target node, the first connected domain is directly determined as the target connected domain.

8. The method according to claim 1, classifying the multiple kernel functions to obtain classification results, comprising: Obtaining a target identifier corresponding to each kernel function, wherein the target identifier is used to indicate whether each kernel function supports the target algorithm; The multiple kernel functions are classified based on the target identifier to obtain the classification result.

9. A model processing device comprising: an acquisition module, configured to acquire a calculation order and a target algorithm of a plurality of kernel functions in an initial model, wherein the plurality of kernel functions are used to represent a plurality of calculation operations of the initial model; a classification module, configured to classify the multiple kernel functions to obtain a classification result, wherein the classification result is used to indicate whether the multiple kernel functions support the target algorithm; a determination module, configured to determine a target subgraph based on the classification result and the calculation order, wherein the target subgraph is used to represent the operation process of a target kernel function, and the target kernel function is used to represent a kernel function that supports the target algorithm; An optimization module, configured to optimize the initial model using the target algorithm and the target subgraph to obtain a target model, wherein the target model is a speech processing model or an image processing model; The device further includes: a first processing model for processing the target speech using the speech processing model to obtain a speech processing result; or a second processing model for processing the target image using the image processing model to obtain an image processing result; The determination module includes: a first determination unit, configured to determine a target directed graph based on the classification result and the calculation order, wherein the target directed graph includes a plurality of nodes and at least one directed edge, the plurality of nodes corresponding to the plurality of kernel functions, and the at least one directed edge being used to represent the calculation order of the plurality of kernel functions; a processing unit, configured to process an initial connected domain of the target directed graph based on a first node to obtain a target connected domain, wherein the kernel function corresponding to the first node does not support the target algorithm, and the target connected domain includes nodes corresponding to the target kernel function; a second determination unit, configured to determine the target subgraph based on the target connected domain; The processing unit includes: a first determining subunit, used to determine the first directed edge corresponding to the first node in response to the fact that the initial connected domain contains the first node and a directed edge exists at the first node; a deleting subunit, used to delete the first directed edge in the initial connected domain to obtain a first connected domain, wherein the first connected domain contains a second node and the kernel function corresponding to the second node supports the target algorithm; and a processing subunit, used to process the first connected domain based on the first node to obtain the target connected domain.

10. The device according to claim 9, characterized in that The optimization modules include: A first determining unit is configured to determine a target operation corresponding to the target algorithm, wherein the target operation is used to represent an operation corresponding to when the target algorithm optimizes the target subgraph; An optimization unit is used to optimize the initial model based on the target operation and the target subgraph to obtain the target model.

11. The device according to claim 10, characterized in that When the target algorithm is a static graph optimization algorithm, the optimization unit includes: an encapsulation subunit, configured to encapsulate the target subgraph into a target static graph; A calling subunit, configured to call the target static graph using a target interface to obtain a calling result; The first optimization subunit is configured to optimize the initial model based on the call result and the target subgraph to obtain the target model.

12. The device according to claim 10, characterized in that When the target algorithm is a close-packed optimization algorithm, the optimization unit also includes: an acquiring subunit, configured to acquire an initial input character, wherein the initial input character is an input character to be input into the target subgraph; a deletion subunit, configured to delete the target symbol from the initial input character to obtain the target input character; an input subunit, configured to input the target input character into the target subgraph and output an initial output character; an adding subunit, configured to add the target symbol to the initial output character to obtain a target output character; The second optimization subunit is configured to optimize the initial model based on the target input character, the target output character, and the target subgraph to obtain the target model.

13. The device according to claim 9, characterized in that The processing unit also includes: The second determining subunit is configured to directly determine the initial connected domain as the target connected domain in response to the initial connected domain not including the first node.

14. The device according to claim 9, characterized in that The processing subunit is further configured to, in response to the presence of a target node in the first connected domain, update the target node in the first connected domain to the first node, to obtain a second connected domain, wherein the target node corresponds to the node pointed to by the first directed edge in the initial connected domain; determine the first directed edge corresponding to the first node; and delete the first directed edge in the second connected domain to obtain the target connected domain.

15. The device according to claim 14, characterized in that The processing subunit is further configured to, in response to the first connected domain not including a target node, directly determine the first connected domain as the target connected domain.

16. The device according to claim 9, characterized in that The classification modules include: an acquiring unit, configured to acquire a target identifier corresponding to each kernel function, wherein the target identifier is used to indicate whether each kernel function supports the target algorithm; A classification unit is used to classify the multiple kernel functions based on the target identifier to obtain the classification result.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Neural network optimization method and device, computer equipment and storage medium

    CN110659728A

  • Graph compiling method and device for calculation graph, equipment and storage medium

    CN111338635A