Neural network model splitting method and apparatus

By matching the neural network model with the structural template in the knowledge base, obtaining the slicing strategy and performing slicing, the problem of low accuracy of slicing strategy in the existing technology is solved, and the performance and computing efficiency of the neural network model are improved.

WO2025091896A1PCT designated stage expired Publication Date: 2025-05-08HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/096737
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-05-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

When the prior art automatically generates the slicing strategy of neural network models, due to the difficulty of performance simulation, the accuracy of the slicing strategy is low, which affects the performance of the computing node.

Method used

By matching the neural network model to be segmented with multiple structural templates in the knowledge base, the slicing strategy of the target structure template is obtained and the target model structure is segmented based on this strategy.

Benefits of technology

The performance of the neural network model is improved, the pressure of cached computing node data to the cache is reduced, and automated segmentation is realized, reducing the workload and debugging difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024096737_08052025_PF_FP_ABST
    Figure CN2024096737_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of neural network models, and discloses a neural network model splitting method and apparatus. The method comprises: matching a neural network model to be split with a plurality of structure templates in a knowledge base, so as to obtain a target structure template matched with a target model structure in the neural network model, wherein the knowledge base further records splitting strategies of the structure templates, the splitting strategies indicate modes of splitting the structure templates, the splitting strategies recorded in the knowledge base can all achieve optimization benefits, and both the model structures and the structure templates comprise a plurality of computing nodes having a specified connection relationship; acquiring a splitting strategy of the target structure template; and on the basis of the splitting strategy of the target structure template, splitting the target model structure to obtain a split neural network model. According to the present application, once the neural network model is split, the pressure on caching data of the computing nodes to a cache can be reduced without affecting the performance of the computing nodes in the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Neural network model segmentation method and device

[0001] This application claims priority to Chinese patent application No. 202311435163.8, filed on October 30, 2023, entitled “Method and Device for Segmenting Neural Network Models,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the technical field of neural network models, and in particular to a method and device for segmenting a neural network model. Background Art

[0003] Typically, data partitioning of the compute nodes of a neural network model can reduce the size of each data copy, allowing the data from the partitioned compute nodes to be cached in a high-speed cache, improving data transfer efficiency. Data partitioning of compute nodes involves partitioning the data of the compute nodes using mathematical equivalence, such that the combination of the multiple copies of data after partitioning is mathematically equivalent to the data before partitioning.

[0004] Currently, computing devices typically automatically generate partitioning strategies using theoretical evaluation methods. In this approach, the computing device needs to evaluate the cost of partitioning, determine the partitioning strategy based on the cost, and automatically generate the number of partitions for each computing node.

[0005] However, due to the great difficulty in simulating the performance of computing nodes, there may be a gap between the results of theoretical evaluation and the performance of actual computing nodes, resulting in the problem of low accuracy in this method of determining the splitting strategy.

[0006] Summary of the Invention

[0007] This application provides a method and apparatus for segmenting a neural network model. This application can improve the performance of a neural network model. For example, after segmenting a neural network model, this application can reduce the pressure of caching data from computing nodes into a high-speed cache without affecting the performance of computing nodes in the neural network model. The technical solutions provided by this application are as follows:

[0008] In a first aspect, the present application provides a method for segmenting a neural network model. The method comprises: matching the neural network model to be segmented with multiple structural templates in a knowledge base to obtain a target structural template that matches the target model structure in the neural network model, wherein the knowledge base further records segmentation strategies for each structural template, wherein the segmentation strategies indicate the manner in which the structural template is segmented, wherein the segmentation strategies recorded in the knowledge base are all capable of achieving optimized benefits, and wherein both the model structure and the structural template include multiple computing nodes with specified connection relationships; obtaining the segmentation strategy of the target structural template; and segmenting the target model structure based on the segmentation strategy of the target structural template to obtain a segmented neural network model.

[0009] In this segmentation method, since the segmentation strategies recorded in the knowledge base can all achieve optimization benefits, when the target model structure in the neural network model matches the target structure template, the target model structure is segmented based on the segmentation strategy of the target structure template, ensuring that the segmented target model structure can achieve optimization benefits compared to the structure before segmentation, so that segmenting the neural network model can achieve positive benefits, thereby improving the performance of the neural network model. For example, after the neural network model is segmented, the pressure of caching the data of the computing nodes into the cache can be reduced without affecting the performance of the computing nodes in the neural network model. In addition, the segmentation process can be automated, which can reduce the workload and debugging difficulty of segmenting the neural network model.

[0010] In one implementation, matching the target model structure with the target structure template includes: structural features of the target model structure are identical to structural features of the target structure template.

[0011] Optionally, the structural features include one or more of the following: topological features and connection features, where the topological features indicate the execution order between computing nodes, and the connection features indicate the connection relationship between computing nodes.

[0012] When multiple target model structures in a neural network model are matched with multiple target structure templates in a knowledge base, there may be model structures containing the same computational nodes among the multiple target model structures. If there are model structures containing the same computational nodes among the multiple target model structures, since a computational node can only be split according to one splitting strategy, further processing is required to ensure that no model structures containing the same computational nodes are present among the multiple split target model structures. Based on the splitting strategy of the target structure template, the target model structure is split, including: when no model structures containing the same computational nodes are present among the multiple target model structures, splitting the target model structures that match the target structure template based on the splitting strategies of the multiple target structure templates; when there are model structures containing the same computational nodes among the multiple target model structures, screening the multiple target model structures so that no model structures containing the same computational nodes are present among the multiple filtered target model structures, and then splitting the filtered target model structures based on the splitting strategies of the target structure templates that match the filtered target model structures.

[0013] In one implementation, multiple target model structures are screened so that no model structure including the same computing node exists in the screened multiple target model structures, including: combining multiple target model structures to obtain multiple model structure combinations, the model structure combinations including some target model structures in the multiple target model structures, and the model structure set does not contain repeated computing nodes; based on the optimization benefits of the segmentation strategies of the multiple target structure templates, respectively obtaining benefits of the multiple model structure combinations; based on the benefits of the multiple model structure combinations, determining a target model structure combination in the multiple model structure combinations, the target model structures in the target model structure combination being the screened target model structures.

[0014] Optionally, before matching the neural network model to be segmented with multiple structural templates, the computing device may pre-split the neural network model into multiple model structures, and then match the multiple model structures with multiple structural templates in the knowledge base. Since the segmentation method of the neural network model provided in this application mainly utilizes the segmentation strategy in the knowledge base to segment the neural network model, the computing device may optionally split the neural network model based on the structural templates in the knowledge base.

[0015] In one implementation, a neural network model is split based on the computational nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model, including: deleting, from all computational nodes in the neural network model, computational nodes that do not conform to the structural characteristics of all structural templates in the knowledge base to obtain the multiple model structures. When the structural characteristics of any computational node in the neural network model are different from the structural characteristics of all computational nodes in the knowledge base, it indicates that there is no computational node in the knowledge base that may match the computational node, and there is no need to match the computational node with the computational nodes in the knowledge base, so the computational node can be deleted from the neural network model.

[0016] In another implementation, the computing device obtains the computing nodes of all structural templates in the knowledge base, and compares the computing nodes in the neural network model with the computing nodes of all structural templates. Then, among all the computing nodes in the neural network model, the computing nodes that are different from all the computing nodes in the knowledge base are deleted to obtain multiple model structures. Since the neural network model is segmented by dividing the data of the computing nodes in a mathematically equivalent manner, and the computing nodes are used to represent a part of the calculation implemented by the neural network model, the neural network model is actually segmented based on the logic of the calculation implemented by the computing nodes. Then, when the computing logic of any computing node in the neural network model is different from the computing logic of all computing nodes in the knowledge base, it means that there is no computing node in the knowledge base that may match the computing node. There is no need to match the computing node with the computing nodes in the knowledge base, and the computing node can be deleted from the neural network model.

[0017] Optionally, before matching the multiple model structures with the structural templates in the knowledge base, the computing device may also screen the multiple model structures, and then match the screened model structures with the multiple structural templates in the knowledge base. By screening the model structures, the total number of model features that need to be matched with the structural templates can be reduced, thereby reducing the total time spent on matching the neural network model to be segmented with the structural templates in the knowledge base, and speeding up the segmentation of the neural network model. In one implementation, the method further includes: screening the multiple model structures based on the number of computing nodes included in the structural templates in the knowledge base. Matching the multiple model structures with the structural templates in the knowledge base includes: matching the screened model structures with the structural templates in the knowledge base.

[0018] In the present application, before matching the neural network model to be segmented with multiple structural templates in the knowledge base, the method also includes: obtaining the structural template based on the segmentation result of the template neural network model; and creating a knowledge base based on the structural template and its segmentation strategy.

[0019] According to the operational logic of the neural network model, data is transferred between computationally related nodes. When a node is split, the data originally provided to it continues to be provided to the multiple nodes it splits from, and the node that originally received data from it continues to receive data from the multiple nodes it splits from. Since all nodes in the same model structure are related, the number of splits between different nodes in the same model structure will be an integer multiple.

[0020] In one implementation, after obtaining multiple model structures, the multiple model structures can be optimized based on the principle of segmentation convergence points and / or segmentation consistent regions, and the optimized model structure is the structural template. The principle based on segmentation convergence points means that the structural template needs to satisfy: all computing nodes except the first computing node (also called the first computing node) and the last computing node in the optimized model structure are segmented. In this way, for the multiple computing nodes in any optimized model structure, except the first computing node and the last computing node, the remaining computing nodes are all computing nodes with segmentation reference significance, so the model structure can be determined as a structural template for reference when segmenting other model structures. Even the first computing node and the last computing node are also segmented.

[0021] The principle of consistent partitioning requires that the optimized model structure meet the following requirements: the number of partitions between different computational nodes in the optimized model structure is an integer multiple. According to the operational logic of the neural network model, data is transferred between computationally related computational nodes. When a computational node is partitioned, the data originally provided to the node will continue to be provided to the multiple computational nodes derived from it, and the computational nodes that originally received data from the node will continue to receive data from the multiple computational nodes derived from it. Computational nodes within the same model structure are all related, so the number of partitions between different computational nodes within the same model structure will be an integer multiple.

[0022] In a second aspect, the present application provides a neural network model segmentation device. The device includes: a matching module for matching the neural network model to be segmented with multiple structural templates in a knowledge base to obtain a target structural template that matches the target model structure in the neural network model, the knowledge base also recording the segmentation strategy of each structural template, the segmentation strategy indicating the method for segmenting the structural template, the segmentation strategy recorded in the knowledge base can all achieve optimization benefits, and the model structure and structural template both include multiple computing nodes with specified connection relationships; a first acquisition module for obtaining the segmentation strategy of the target structural template; a segmentation module for segmenting the target model structure based on the segmentation strategy of the target structural template to obtain a segmented neural network model.

[0023] Optionally, matching the target model structure with the target structure template includes: structural features of the target model structure are identical to structural features of the target structure template.

[0024] Optionally, the structural features include one or more of the following: topological features and connection features, where the topological features indicate the execution order between computing nodes, and the connection features indicate the connection relationship between computing nodes.

[0025] Optionally, when multiple target model structures in the neural network model are respectively matched with multiple target structure templates in the knowledge base, the segmentation module is specifically used to: when there is no model structure including the same computing node among the multiple target model structures, segment the target model structures that match the target structure template based on the segmentation strategies of the multiple target structure templates; when there is a model structure including the same computing node among the multiple target model structures, filter the multiple target model structures so that there is no model structure including the same computing node among the multiple target model structures that have been filtered, and segment the filtered target model structures based on the segmentation strategies of the target structure templates that match the filtered target model structures.

[0026] Optionally, the segmentation module is specifically used to: combine multiple target model structures to obtain multiple model structure combinations, where the model structure combination includes some target model structures in the multiple target model structures, and the model structure set does not contain repeated computing nodes; based on the optimization benefits of the segmentation strategy of multiple target structure templates, respectively obtain the benefits of the multiple model structure combinations; based on the benefits of the multiple model structure combinations, determine the target model structure combination in the multiple model structure combinations, and the target model structure in the target model structure combination is the screened target model structure.

[0027] Optionally, the matching module is specifically used to: split the neural network model based on the computing nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model; and match the multiple model structures with the structural template in the knowledge base.

[0028] Optionally, the matching module is specifically used to: delete the computing nodes that do not conform to the structural features of all the structural templates in the knowledge base among all the computing nodes of the neural network model, and obtain multiple model structures.

[0029] Optionally, the matching module is specifically used to: screen multiple model structures based on the number of computing nodes included in the structure template in the knowledge base; and match the screened model structures with the structure template in the knowledge base.

[0030] Optionally, the apparatus further includes: a second acquisition module and a creation module. The second acquisition module is configured to acquire a structure template based on the segmentation result of the template neural network model. The creation module is configured to create a knowledge base based on the structure template and its segmentation strategy.

[0031] Optionally, the structural template satisfies one or more of the following: the number of shares of different computing nodes in the structural template is an integer multiple; and all computing nodes except the first computing node and the last computing node in the structural template are split.

[0032] Optionally, both the first compute node and the last compute node are split.

[0033] In a third aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores program instructions, and the processor runs the program instructions to execute the method provided in the first aspect of the present application and any possible implementation thereof.

[0034] In a fourth aspect, the present application provides a computing device cluster, comprising multiple computing devices, wherein the multiple computing devices include multiple processors and multiple memories, wherein program instructions are stored in the multiple memories, and the multiple processors execute the program instructions, so that the computing device cluster executes the method provided in the first aspect of the present application and any possible implementation thereof.

[0035] In a fifth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes program instructions. When the program instructions are executed on a computing device, the computing device executes the method provided in the first aspect of the present application and any possible implementation thereof.

[0036] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when run on a computer, enables the computer to execute the method provided in the first aspect of the present application and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a schematic diagram of a computational graph of a neural network model provided in an embodiment of the present application;

[0038] FIG2 is a schematic diagram of a computation graph after segmenting the exponential operation node in FIG1 , provided by an embodiment of the present application;

[0039] FIG3 is a schematic structural diagram of an implementation scenario involved in a segmentation method of a neural network model provided in an embodiment of the present application;

[0040] FIG4 is a schematic structural diagram of an implementation scenario involved in another neural network model segmentation method provided in an embodiment of the present application;

[0041] FIG5 is a flow chart of a segmentation method of a neural network model provided in an embodiment of the present application;

[0042] FIG6 is a schematic diagram of a process for matching a model structure with a structural template according to an embodiment of the present application;

[0043] FIG7 is a flow chart of a segmentation strategy based on a target structure template provided in an embodiment of the present application for segmenting a target model structure;

[0044] FIG8 is a schematic diagram of a computational graph of a neural network model to be segmented provided in an embodiment of the present application;

[0045] FIG9 is a schematic diagram of a structural template provided in an embodiment of the present application;

[0046] FIG10 is a flowchart of a method for screening multiple target model structures so that no model structure including the same computing node exists among the screened target model structures, provided by an embodiment of the present application;

[0047] FIG11 is a schematic diagram of a multiple target model structure including computing nodes and benefits provided in an embodiment of the present application;

[0048] FIG12 is a schematic diagram of a deployment method of a neural network model provided in an embodiment of the present application;

[0049] FIG13 is a schematic diagram of another deployment method of a neural network model provided in an embodiment of the present application;

[0050] FIG14 is a flowchart of creating a knowledge base provided by an embodiment of the present application;

[0051] FIG15 is a schematic diagram of a process for obtaining a structural template provided in an embodiment of the present application;

[0052] FIG16 is a schematic diagram of a process of a segmentation method of a neural network model provided in an embodiment of the present application;

[0053] FIG17 is a schematic diagram of functional modules of a method for segmenting a neural network model provided in an embodiment of the present application;

[0054] FIG18 is a schematic diagram of a segmentation device for a neural network model provided in an embodiment of the present application;

[0055] FIG19 is a schematic diagram of another neural network model segmentation device provided in an embodiment of the present application;

[0056] FIG20 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0057] FIG21 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0059] To facilitate understanding, the technology and background involved in the embodiments of this application are explained below.

[0060] A neural network model typically includes multiple connected neuron nodes. Each neuron node is responsible for one or more computational operations, such as matrix multiplication and addition. A neural network model can be represented by multiple computational nodes. The functionality of each computational node can be implemented by one or more neuron nodes. A computational node can be considered a collection of one or more operations performed on an operation object. A computational node is also called an operator. When a neural network model is run on a computing device, the data required for computational node operations needs to be cached in memory. For example, when a neural network model is run on a chip, the data required for computational operations of each computational node in the neural network model needs to be cached in memory. However, when the amount of input data to a computational node is large, the input data cannot be cached in memory at high speed, causing the process of moving the data to the cache to become a performance bottleneck for the operation of the neural network model, affecting the computational efficiency of the neural network model.

[0061] Typically, data partitioning of the compute nodes of a neural network model can reduce the size of each data copy, allowing the data from the partitioned compute nodes to be cached in a high-speed cache, improving data transfer efficiency. Data partitioning of compute nodes involves partitioning the data of the compute nodes using mathematical equivalence, such that the combination of the multiple copies of data after partitioning is mathematically equivalent to the data before partitioning.

[0062] For example, Figure 1 is an example of a computation graph of a neural network model provided in an embodiment of the present application. As shown in Figure 1, the neural network model includes sequentially connected data nodes (such as Data in Figure 1), exponential operation nodes (such as Exp in Figure 1) and output nodes (such as Output in Figure 1). The data node, the exponential operation node and the output node are all computation nodes. The data node is used to input data to the exponential operation node. The exponential operation node is used to obtain the exponential of the input data with the natural constant e as the base. The output node is used to output the calculation result of the exponential operation node. Figure 2 is a computation graph provided in an embodiment of the present application after the exponential operation node in Figure 1 is split. As shown in Figure 2, the exponential operation node is split into two, and the input end of each exponential operation node after the split is connected to the output end of the data node, and the output end of the exponential operation node after the split is connected to the input end of the output node. It can be seen that the combination of the input data of the two exponential operation nodes after the split is mathematically equivalent to the input data of the exponential operation node before the split.

[0063] Among them, the computational graph is a form of representation of the neural network model. The computational graph usually includes multiple graph nodes with a connection relationship. The graph nodes are used to represent the computational nodes in the neural network model. The edges between the graph nodes are used to represent the connection relationship between the computational nodes in the neural network model. When there is a connection relationship between the computational nodes, it means that there is data transfer between the two computational nodes. It should be noted that the use of a computational graph to represent the neural network model here is only an example. The neural network model can also be represented by other types of graphs, and the embodiments of the present application do not specifically limit it.

[0064] Currently, there are two main methods for determining the splitting strategy for compute nodes in neural network models. One is for manual determination of the splitting strategy based on expert strategies. The other is for computing devices to automatically generate the splitting strategy using theoretical evaluation and other methods. The first method requires manual analysis based on each neural network model and manual modification of the neural network model script file to split and merge nodes in the neural network model to ensure that the compute nodes are mathematically equivalent before and after the split. For example, for a computational graph, a human can manually specify the number of splits for each graph node and manually control the order of each graph node in the computational graph to reduce the amount of data and cache the data. However, manually specifying each graph node is labor-intensive, cumbersome, and error-prone. It also requires frequent modifications to the network model structure, making debugging difficult.

[0065] In the second method of determining the splitting strategy, the computing device needs to evaluate the cost of splitting, determine the splitting strategy based on the splitting cost, and automatically generate the number of splits for each computing node. However, due to the great difficulty in simulating the performance of computing nodes, there may be a gap between the results of theoretical evaluation and the performance of actual computing nodes. There is no way to guarantee that all neural network models can achieve positive returns compared to before splitting. Some neural network models may even have negative returns due to inaccurate evaluation. At the same time, it is also difficult to simulate the behavior of cache, resulting in low accuracy in this method of determining the splitting strategy. Moreover, this impact is often not linear for the performance of computing nodes of various sizes.

[0066] As can be seen from the above, data segmentation of neural network models may cause a decrease in computing node performance. Therefore, how to reasonably segment neural network models is a relatively difficult problem to solve.

[0067] To this end, an embodiment of the present application provides a method for segmenting a neural network model. In this method, the computing device first matches the neural network model to be segmented with multiple structural templates in the knowledge base to obtain a target structural template that matches the target model structure in the neural network model; then, the segmentation strategy of the target structural template is obtained; and then, based on the segmentation strategy of the target structural template, the target model structure is segmented to obtain a segmented neural network model. Among them, both the model structure and the structural template include multiple computing nodes with specified connection relationships. The knowledge base records multiple structural templates and the segmentation strategy of each structural template. The segmentation strategy indicates the way to segment the structural template, and the segmentation strategies recorded in the knowledge base can all achieve optimization benefits. Obtaining optimization benefits here means obtaining positive benefits.

[0068] Since the segmentation strategies recorded in the knowledge base can all achieve optimization benefits, when the target model structure in the neural network model matches the target structure template, the target model structure is segmented based on the segmentation strategy of the target structure template, ensuring that the segmented target model structure can achieve optimization benefits compared to the structure before segmentation, so that segmentation of the neural network model can achieve positive benefits, thereby improving the performance of the neural network model. For example, after segmenting the neural network model, it is possible to reduce the pressure of caching the data of the computing nodes to the cache without affecting the performance of the computing nodes in the neural network model. In addition, the segmentation process can be automated, which can reduce the workload and debugging difficulty of segmenting the neural network model.

[0069] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, method flow, hardware devices, and software devices.

[0070] The following first illustrates the application scenarios of the embodiments of the present application.

[0071] FIG3 is a schematic diagram of a structure of an implementation scenario involving a segmentation method for a neural network model provided in an embodiment of the present application. As shown in FIG3 , the implementation scenario includes a computing device 10. By executing the segmentation method for a neural network model provided in an embodiment of the present application, the computing device 10 can segment the neural network model.

[0072] In one implementation, the neural network model segmentation method provided in the embodiment of the present application can be implemented by running an executable program on the computing device 10. For example, the computing device 10 is configured with a processor and a memory, and the memory stores an executable program for implementing the neural network model segmentation method. The processor executes the executable program, and the computing device 10 can execute the neural network model segmentation method. The processor can be a neural network processing unit (NPU) or a graphics processing unit (GPU). In addition, the executable program of the neural network model segmentation method can also be presented in the form of an application installation package. After the application installation package is installed in the computing device 10, the neural network model segmentation method can be implemented by running the executable program.

[0073] Optionally, the computing device 10 may be a terminal. The terminal may be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, a smart vehicle-mounted device, or an Internet of Things device.

[0074] FIG4 is a schematic diagram of the structure of an implementation scenario involved in another neural network model segmentation method provided in an embodiment of the present application. As shown in FIG4 , the implementation scenario may include a computing device 10 and a client 20. The client 20 is capable of establishing a communication connection with the computing device 10. For example, a communication connection can be established between the client 20 and the computing device 10 via a network. Optionally, the network can be a local area network, the Internet, or other networks, which are not limited in the embodiments of the present application.

[0075] In the implementation scenario of FIG4 , the user can interact with the computing device 10 through the client 20. For example, the user can provide the computing device 10 with a segmentation instruction and a model file of the neural network model to be segmented through the client 20. In one implementation, the segmentation instruction can optionally carry the model file of the neural network model to be segmented. Accordingly, the computing device 10 can execute the segmentation method of the neural network model provided in the embodiment of the present application based on the segmentation instruction of the client 20 to segment the neural network model.

[0076] In one implementation, the client 20 can be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, a smart vehicle-mounted device or an Internet of Things device. The computing device 10 can be a server, or a server cluster consisting of several servers, or a cloud computing service center. Among them, a large number of basic resources owned by the cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources and network resources are deployed in the cloud computing service center. The cloud computing service center can use this large amount of basic resources to implement the segmentation method of the neural network model provided in the embodiment of the present application.

[0077] When the computing device 10 is implemented through a cloud computing service center, the user can access the cloud platform through the client 20 and provide the computing device 10 with a segmentation instruction and a neural network model to be segmented through the cloud platform. At this time, the computing device 10 in the cloud platform can provide the user with the function of automatically segmenting the neural network model. Optionally, this function can be abstracted into a segmentation cloud service by the cloud service provider on the cloud platform. After the user purchases the segmentation cloud service on the cloud platform, the cloud platform can use the resources of the cloud computing service center and use the segmentation method of the neural network model provided in the embodiment of the present application to segment the neural network model owned by the user. In addition, the segmentation cloud service can be provided as an independent cloud service or as an additional service of other cloud services. In addition, the cloud platform can be a cloud platform of a central cloud, a cloud platform of an edge cloud, or a cloud platform including a central cloud and an edge cloud, which is not specifically limited in the embodiment of the present application. It should be noted that in the implementation scenario shown in Figure 4, the computing device 10 can also be implemented through other resource platforms besides the cloud platform, which is not specifically limited in the embodiment of the present application.

[0078] It should be understood that the above content is an exemplary description of the implementation scenario of the segmentation method of the neural network model provided in the embodiment of the present application, and does not constitute a limitation on the implementation scenario of the segmentation method of the neural network model. For example, the idea of ​​solving technical problems by the segmentation method of the neural network model provided in the embodiment of the present application can also be applied to other fields related to graphs. For example, it is applied to the field of chemistry, and the knowledge base of the chemical field records the chemical molecular structure. Based on the idea of ​​the present application, the chemical structure to be identified can be matched with the chemical molecular structure recorded in the knowledge base, and then the identification of the chemical structure can be realized based on the chemical molecular structure obtained by matching, thereby accelerating the identification of chemical molecules, etc. It is known to those skilled in the art that as business needs change, its implementation scenario can be adjusted according to application requirements, and the embodiments of the present application do not list them one by one.

[0079] The following describes a segmentation method for a neural network model provided in an embodiment of the present application. The method can be executed by the computing devices shown in Figures 3 and 4. As shown in Figure 5, the segmentation method for a neural network model includes the following steps:

[0080] Step 501: Split the neural network model based on the computing nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model. The knowledge base records multiple structural templates and the segmentation strategy of each structural template. The segmentation strategy indicates the way to segment the structural template. The segmentation strategies recorded in the knowledge base can all achieve optimization benefits. The model structure and the structural template both include multiple computing nodes with specified connection relationships.

[0081] Before matching the neural network model to be segmented with multiple structural templates, the computing device may optionally pre-split the neural network model into multiple model structures, and then match the multiple model structures with multiple structural templates. In the segmentation method of the neural network model provided in this application, the segmentation strategy in the knowledge base is mainly used to segment the neural network model. Therefore, the computing device may optionally split the neural network model based on the structural template in the knowledge base. The following two splitting methods are used as examples to illustrate the implementation method of the computing device splitting the neural network model based on the structural template:

[0082] In a first implementation, a computing device obtains the computational nodes of all structural templates in a knowledge base, compares the computational nodes in the neural network model with the computational nodes of all structural templates, and then deletes the computational nodes in the neural network model that do not conform to the structural characteristics of all structural templates in the knowledge base, thereby obtaining multiple model structures. Because all computational nodes in the neural network model are organized into a single network, deleting one or more of these computational nodes may cause the neural network model to be divided into at least two subnets. These subnets are the resulting model structures.

[0083] Structural features include one or more of the following: topological features (also known as topological order features) and connection features. Topological features indicate the execution order between computing nodes. For example, the topological features of the model structure shown in Figure 1 indicate that the data node, the exponential operation node, and the output node are executed in sequence. Connection features indicate the connection relationship between computing nodes. For example, the connection features of the model structure shown in Figure 1 indicate that the data node inputs data to the exponential operation node, the exponential operation node inputs data to the output node, and the output node outputs the output result of the entire model structure. When the structural features include connection features, the first implementation method is equivalent to comparing the connection method of all computing nodes in the neural network model with the connection method of all computing nodes in the knowledge base, and then deleting the computing nodes in the neural network model whose connection method is different from the connection method of all computing nodes in the knowledge base, so as to split the neural network model into multiple model structures. When the structural features include topological features, the first implementation method is equivalent to comparing the execution order of all computing nodes in the neural network model with the execution order of all computing nodes in the knowledge base, and then deleting the computing nodes in the neural network model whose execution order is different from the execution order of all computing nodes in the knowledge base, so as to split the neural network model into multiple model structures. When the structural characteristics of any computing node in the neural network model are different from the structural characteristics of all computing nodes in the knowledge base, it means that there is no computing node in the knowledge base that may match the computing node. There is no need to match the computing node with the computing nodes in the knowledge base, so the computing node can be deleted from the neural network model.

[0084] In a second implementation, a computing device obtains the computational nodes of all structural templates in a knowledge base, compares the computational nodes in a neural network model with the computational nodes of all structural templates, and then deletes the computational nodes that are different from all computational nodes in the knowledge base from all computational nodes in the neural network model, thereby obtaining multiple model structures. The computational nodes in the neural network model are different from those in the knowledge base, and optionally, the computational logic implemented by the computational nodes in the neural network model and the computational nodes in the knowledge base are different. Since segmenting a neural network model involves segmenting the data of the computational nodes in a mathematically equivalent manner, and the computational nodes are used to represent a portion of the computation implemented by the neural network model, segmenting the neural network model is actually segmenting based on the logic of the computation implemented by the computational nodes. If the computational logic of any computational node in the neural network model is different from the computational logic of all computational nodes in the knowledge base, it indicates that there is no computational node in the knowledge base that can match the computational node. Therefore, there is no need to match the computational node with the computational nodes in the knowledge base, and the computational node can be deleted from the neural network model. For example, as shown in Figure 6, the computing device first obtains the computing nodes of all structural templates in the knowledge base to obtain a computing node set. Then, the computing device compares the computing nodes in the neural network model to be split with all computing nodes in the computing node set to determine whether there are computing nodes in the computing node set that are identical to the computing nodes in the neural network model to be split. Then, among all the computing nodes in the neural network model to be split, the computing nodes that are different from all computing nodes in the knowledge base are deleted, thereby obtaining multiple model structures.

[0085] It should be noted that the above implementation method of measuring whether the computing nodes in the neural network model are the same as the computing nodes in the knowledge base is only an example. Whether the computing nodes in the neural network model are the same as the computing nodes in the knowledge base can also be measured by other methods. The other methods can be optionally decided based on application requirements, and the embodiments of this application do not list them one by one.

[0086] Optionally, after splitting the neural network model into multiple model structures, the computing device may further filter the multiple model structures and then match the filtered model structures with multiple structural templates in the knowledge base. By filtering the model structures, the total number of model features that need to be matched with the structural templates can be reduced, thereby reducing the total time spent matching the neural network model to be split with the structural templates in the knowledge base, and accelerating the segmentation of the neural network model. In one implementation, the computing device may optionally filter the multiple model structures based on the number of computational nodes included in the structural templates in the knowledge base. For example, as shown in Figure 6, the computing device obtains the minimum total number of computational nodes included in the structural templates in the knowledge base, and obtains the total number of computational nodes included in each model structure. From the multiple model structures, the computing device deletes the model structures whose total number of computational nodes is less than the minimum value. The remaining model structures are the filtered model structures. If the total number of computational nodes included in the model structure is less than the minimum value, it indicates that there is no structural template in the knowledge base that can potentially match the model structure. There is no need to match the model structure with the structural templates in the knowledge base, and the model structure can be deleted.

[0087] For another example, the computing device obtains the total number of computing nodes included in all structural templates in the knowledge base, and obtains the total number of computing nodes included in each model structure. Among multiple model structures, the computing device deletes the model structures whose total number of computing nodes is different from the total number of computing nodes included in any structural template in the knowledge base. The remaining model structures are the filtered model structures. If the total number of computing nodes included in the model structure is different from the total number of computing nodes included in any structural template in the knowledge base, it indicates that there is no structural template in the knowledge base that can potentially match the model structure. There is no need to match the model structure with the structural templates in the knowledge base, and the model structure can be deleted.

[0088] Furthermore, before operating on the neural network model, the neural network model can optionally be converted into a computational graph, and the structural templates in the knowledge base can optionally also be represented using computational graphs. In this case, the processing of the neural network model based on the structural template is computational graph-based processing. For example, as shown in Figure 6, the computing device first obtains the computational graph of each structural template in all structural templates in the knowledge base (hereinafter referred to as the template computational graph), obtains the computational nodes of the structural template based on the template computational graph, thereby obtaining a computational node set containing the computational nodes in all structural templates, and obtains the computational graph of the neural network model to be split, obtaining all the computational nodes of the neural network model to be split. Then, the computing device compares the computational nodes in the neural network model to be split with all the computational nodes in the computational node set. The computational graph can optionally be a directed acyclic graph (DAG). In one implementation, the computing device can optionally convert the neural network model into a computational graph based on the in-degree of the computational nodes. For example, the computing device obtains the in-degree of each computational node, and then divides the computational nodes with the same in-degree into the same layer, resulting in multiple layers, where the computational nodes contained in the multiple layers have different in-degrees, and the computational nodes contained in the same layer have the same in-degree. Then, according to the size of the in-degree of the computing nodes contained in the layer, the layers containing computing nodes with smaller in-degree are sorted before the layers containing computing nodes with larger in-degree. The resulting set including multiple layers with a sequence between the layers is the computation graph. Here, in-degree refers to the sum of the number of times a point in a directed graph serves as the end point of an edge in the graph. It should be understood that the neural network model here is represented by a computation graph, and the implementation method of obtaining the computation graph is an example. The neural network model can also be represented by other types of graphs, and the computation graph can also be obtained by other methods. The embodiments of the present application do not specifically limit them.

[0089] It should be noted that step 501 is an optional step. That is, before matching the neural network model with multiple structural templates in the knowledge base, the computing device may optionally perform step 501 to split the neural network model into multiple model structures, and then match these multiple model structures with the multiple structural templates in the knowledge base. Alternatively, the computing device may not perform step 501, that is, not split the neural network model into multiple model structures, and in this case, may match the entire neural network model with the structural templates in the knowledge base.

[0090] Step 502: Match the neural network model to be segmented with multiple structural templates in the knowledge base to obtain a target structural template that matches the target model structure in the neural network model.

[0091] After the computing device obtains the neural network model, it can match the neural network model with multiple structural templates in the knowledge base to obtain a target structural template that matches the target model structure in the neural network model. The target model structure is any model structure in the neural network model that matches the structural template in the knowledge base. When step 501 is executed, step 502 is to match the multiple model structures of the neural network model to be split with the multiple structural templates in the knowledge base. When step 501 is not executed, step 502 is to match the entire neural network model to be split with the multiple structural templates in the knowledge base. For example, when step 501 splits the neural network model into multiple model structures and screens the multiple model structures, the schematic diagram of the overall process of steps 501 and 502 can be optionally shown as Figure 6. After obtaining multiple model structures according to the steps shown in 6, the computing device continues to match the multiple model structures with the multiple structural templates in the knowledge base. If step 501 is not performed, the same computing node may be divided into multiple model structures according to different division methods. When matching the neural network model with the structural template in the knowledge base, the multiple model structures need to be matched with all the structural templates in the knowledge base, and the matching task volume is large. Moreover, when the scale of the neural network model increases, the time complexity of the matching will increase exponentially. When the aforementioned step 501 is performed, since the objects to be matched in the neural network model are pre-screened before matching the neural network model with the structural template, when matching the neural network model with the structural template in the knowledge base, only the screened model structure needs to be compared with the structural template, which can effectively reduce the time complexity of the matching and thus speed up the matching speed.

[0092] In one embodiment, the target model structure and the target structure template are matched, including: the structural features of the target model structure are the same as the structural features of the target structure template. When the structural features include topological features, step 502 is to compare the execution order between the computing nodes included in the model structure with the execution order between the computing nodes included in the structure template. When the execution order between the computing nodes included in the model structure is the same as the execution order between the computing nodes included in the structure template, it is determined that the model structure matches the structure template. When the structural features include connection features, step 502 is to compare the connection relationship between the computing nodes included in the model structure with the connection relationship between the computing nodes included in the structure template. When the connection relationship between the computing nodes included in the model structure is the same as the connection relationship between the computing nodes included in the structure template, it is determined that the model structure matches the structure template.

[0093] In one implementation, before the computing device matches the model structure with the structure template, it may first use a byte sequence to represent the model structure and the structure template, and then use a byte sequence matching method, such as the KMP algorithm, to match the byte sequence of the model structure with the byte sequence of the structure template.

[0094] In addition, when the neural network model and the structural template in the knowledge base are both represented by a computational graph including multiple layers, when matching the neural network model to be segmented with any structural template, the hierarchical structural features of the computational graph of the neural network model can be optionally compared with the hierarchical structural features of the computational graph of the structural template. When the computational graph of the neural network model has the hierarchical structural features of the computational graph of the structural template, the neural network model and the structural template are matched in the above manner. Otherwise, there is no need to continue matching. The hierarchical structural features include: the total number of multiple layers in the computational graph, the order of each layer in the computational graph, the number of computational nodes included in each layer, and the in-degree of the computational nodes in the layer. Optionally, the hierarchical structural features of the computational graph of the neural network model are compared with the hierarchical structural features of the computational graph of the structural template, including: performing a matching process on each layer of the computational graph of the neural network model in order from small to large layers. The matching process includes: first comparing the computational nodes in the i-th (for example, i=1) layer of the computational graph of the neural network model with all layers of the computational graph of the structural template. When the computational nodes in the i-th layer match the computational nodes in the j-th layer of the computational graph of the structural template, continue to compare the computational nodes in the i+1-th layer of the computational graph of the neural network model with the computational nodes in the j-th layer to the last layer of the computational graph of the structural template, and so on, until the comparison of each layer of the computational graph of the neural network model is completed. If the computational nodes in each layer of the computational graph of the neural network model can find matching points in the computational graph of the structural template, then the structural template is said to match the neural network model, that is, the computational graph of the neural network model to be tested has the hierarchical structural characteristics of the computational graph of the structural template. Otherwise, the computational graph of the neural network model to be tested does not have the hierarchical structural characteristics of the computational graph of the structural template. Among them, the layer where the computational node with the smallest in-degree is located has the smallest hierarchical order.

[0095] By performing a preliminary match on the neural network model based on its hierarchical structure, the matching speed of the neural network model can be accelerated. Furthermore, since the layers in the computational graph have an order, this preliminary match ensures that the execution order of the computational nodes does not violate the execution order of the computational nodes in the structural template, further ensuring the accuracy of the match.

[0096] Step 503: Obtain a segmentation strategy for the target structure template.

[0097] Since the knowledge base records multiple structural templates and the segmentation strategies of each structural template, after obtaining the target structural template that matches the target model structure, the computing device can obtain the segmentation strategy of the target structural template from the knowledge base, so as to segment the target model structure based on the segmentation strategy. Among them, the segmentation strategy indicates the way to segment the structural template. The segmentation strategies recorded in the knowledge base can all achieve optimization benefits. For example, assume that Figure 1 is a schematic diagram of a structural template. As shown in Figure 1, the target structural template includes three computing nodes, and its segmentation strategy is: do not segment the first computing node and the third computing node of the target structural template, and divide the second computing node of the target structural template into two computing nodes. In one implementation, the knowledge base records multiple structural templates, multiple segmentation strategies, and the corresponding relationship between multiple structural templates and segmentation strategies. The computing device can optionally query the corresponding relationship based on the target structural template to obtain the segmentation strategy corresponding to the target structural template.

[0098] Step 504: Based on the segmentation strategy of the target structure template, the target model structure is segmented to obtain a segmented neural network model.

[0099] After the computing device obtains the segmentation strategy of the target structure template, it can segment the target model structure that matches the target structure template based on the segmentation strategy to obtain a segmented neural network model. However, when multiple target model structures in the neural network model are respectively matched with multiple target structure templates in the knowledge base, if there are model structures including the same computing nodes in the multiple target model structures, since a computing node can only be segmented according to one segmentation strategy, further processing is required to ensure that there are no model structures including the same computing nodes in the multiple segmented target model structures. For example, as shown in FIG7 , segmenting the target model structure based on the segmentation strategy of the target structure template includes:

[0100] Step 5041: When there is no model structure including the same computing node among the multiple target model structures, the target model structures matching the target structure template are segmented based on the segmentation strategies of the multiple target structure templates.

[0101] When there is no model structure including the same computing node among multiple target model structures, the computing device can directly segment the target model structure that matches the target structure template according to the segmentation strategy of multiple target structure templates, and then splice the segmented target model structure and the model structure that does not need to be segmented according to the order of each model structure in the neural network model to obtain a segmented neural network model.

[0102] For example, assume that Figure 8 is a computational graph of a neural network model to be split, and the numbers in the circles in Figure 8 are used to identify different computational nodes. As shown in Figure 8, the neural network model includes two computational nodes as inputs, the outputs of the two computational nodes are connected to the same computational node, and four computational nodes are sequentially connected to the computational node. After step 502, it can be obtained that the neural network model includes model structure 1, model structure 2 and a computational node. Model structure 1 matches the structural template shown in Figure 9, and model structure 2 matches the structural template shown in Figure 1. After step 503, the splitting strategy of the structural template shown in Figure 1 is: the first and third computational nodes of the target structural template are not split, and the second computational node of the target structural template is split into two computational nodes. The splitting strategy of the structural template shown in Figure 9 is: one input node in the structural template is split into two computational nodes, and the remaining nodes are not split. Since the model structure 1 and model structure 2 include different computing nodes, the computing device can directly split model structure 1 and model structure 2 according to the segmentation strategy of the structural template shown in Figures 1 and 9, and then splice the last computing node of model structure 1 with the first computing node of model structure 2, and splice the third computing node of model structure 2 with the last computing node of the neural network model to obtain a segmented neural network model.

[0103] Step 5042: When there are model structures including the same computing nodes among multiple target model structures, the multiple target model structures are screened so that there is no model structure including the same computing nodes among the screened multiple target model structures, and then the screened target model structures are segmented based on the segmentation strategy of the target structure template that matches the screened target model structures.

[0104] When there are model structures including the same computing nodes among multiple target model structures, the computing device may optionally first filter the multiple target model structures so that there is no model structure including the same computing nodes among the filtered target model structures, and then split the filtered target model structure according to the splitting strategy that matches the filtered target model structure, and then splice the split target model structure and the model structure that does not need to be split according to the order of each model structure in the neural network model to obtain a split neural network model.

[0105] In one implementation, as shown in FIG10 , the implementation process of screening multiple target model structures so that no model structure including the same computing node exists among the screened target model structures includes:

[0106] Step 5042a: Combine multiple target model structures to obtain multiple model structure combinations, where any model structure combination includes part of the target model structures in the multiple target model structures, and the model structure set does not contain duplicate computing nodes.

[0107] The process of combining multiple target model structures by the computing device is actually to screen the target model structure from all model structures that match the structural template, and ensure that the computing nodes included in each two model structures screened into a model structure combination are different during the screening process. For example, assume that Figure 11 is a schematic diagram of multiple target model structures that match the structural template in the knowledge base, including computing nodes and benefits. Each horizontal line in Figure 11 represents a target model structure, and the horizontal axis covered by the horizontal line represents the sequence number of the computing nodes included in the target model structure. As shown in Figure 11, there are 5 target model structures Q1 to Q5 in the neural network model that match the structural template in the knowledge base. Target model structure Q1 includes computing nodes 1 to 6. Target model structure Q2 includes computing nodes 3 to 9. Target model structure Q3 includes computing nodes 7 to 11. Target model structure Q4 includes computing nodes 5 to 14. Target model structure Q5 includes computing nodes 11 to 16. Based on the principle that any model structure set does not contain duplicate computation nodes, the five target model structures are combined to obtain the following model structure combinations: Model structure combination 1 includes target model structure Q1 and target model structure Q3, model structure combination 2 includes target model structure Q2 and target model structure Q5, and model structure combination 3 includes target model structure Q4. It can be seen that each model structure combination includes some target model structures from multiple target model structures, and each model structure set does not contain duplicate computation nodes.

[0108] Step 5042b: Based on the optimization benefits of the segmentation strategies of the multiple target structure templates, the benefits of the multiple model structure combinations are obtained respectively.

[0109] After obtaining any model structure combination, the computing device obtains the optimization benefits of the target model structures included in the model structure combination, and obtains the benefit of the model structure combination based on the optimization benefits of the multiple target model structures included in the model structure combination. For example, the benefit of the model structure combination can be equal to the sum of the optimization benefits of all target model structures in the model structure combination. The benefit of the target model structure refers to the optimization benefit of the segmentation strategy of the target structure template that matches the target model structure. For example, continuing to illustrate with Figure 11, the vertical axis in Figure 11 represents the optimization benefit that can be achieved by segmenting the target model structure according to the segmentation strategy of the target structure template that matches the target model structure. As shown in Figure 11, the optimization benefits of Q1 to Q5 are 6, 8, 5, 7, and 4, respectively. Since model structure combination 1 includes target model structure Q1 and target model structure Q3, model structure combination 2 includes target model structure Q2 and target model structure Q5, and model structure combination 3 includes target model structure Q4, the benefit of model structure combination 1 is 11, the benefit of model structure combination 2 is 12, and the benefit of model structure combination 3 is 7.

[0110] Step 5042c: Based on the benefits of the multiple model structure combinations, determine a target model structure combination from the multiple model structure combinations, where the target model structure in the target model structure combination is a screened target model structure.

[0111] After the computing device obtains the benefits of multiple model structure combinations, it can select a target model structure combination from the multiple model structure combinations. Since the model structure set does not contain duplicate computing nodes, the target model structure in the target model structure combination is a screened target model structure. In one implementation, the computing device may optionally adopt the principle of maximizing total benefits to select a target model structure combination from multiple model structure combinations. That is, the computing device may determine the model structure combination with the largest benefit among the multiple model structure combinations as the target model structure combination. For example, continuing to use Figure 11 as an example, since the benefit of model structure combination 1 is 11, the benefit of model structure combination 2 is 12, and the benefit of model structure combination 3 is 7, it can be obtained that model structure combination 2 is the target model structure combination, and the target model structure Q2 and target model structure Q5 are screened target model structures.

[0112] The process of determining the target model structure combination is essentially to use dynamic programming to solve the problem of the model structure combination with the highest return matching. It can be expressed using the following dynamic programming equation.

[0113] Where i is the model structure index, and q[i] is the benefit of model structure i. dp[i-1] is the benefit of the model structure combination containing model structure i-1, which was traversed before the current model structure i, when traversing multiple model structures. frt[i] represents the number of the model structure combination with the maximum benefit among the model structure combinations with numbers less than i that do not contain duplicate nodes with model structure i. Thus, dp[frt[i]+q[i]] is the benefit of the model structure combination with the maximum benefit among the model structure combinations that include model structure i. This dynamic programming equation actually compares dp[i-1] and q[i]+dp[frt[i]] for each model structure i, taking into account the state transition equation, and taking the larger value as dp[i]. This state equation is then used to determine the structure combination with the maximum benefit. This shows that determining the target model structure combination through the dynamic programming equation can ensure the maximum benefit of the knowledge base.

[0114] After the neural network model is segmented, the segmented neural network model can be deployed on a computing device for implementation. For example, the segmented neural network model can be deployed on a chip for implementation. The deployment options include at least the following two: As shown in Figure 12, the segmented neural network model can be deployed on a chip with a cache. In this case, the cache is used to cache data during the calculation process for the chip. As shown in Figure 13, the segmented neural network model is deployed on multiple chips, and the multiple computing nodes that have been segmented are deployed on multiple chips respectively.

[0115] The above-mentioned process of segmenting the neural network model based on the knowledge base is the application process of the knowledge base. In contrast to the application process of the knowledge base, the embodiment of the present application also provides a process for creating a knowledge base. As shown in Figure 14, the process of creating a knowledge base includes the following steps:

[0116] Step 1401: Obtain a structural template based on the segmentation result of the template neural network model.

[0117] A knowledge base can be created based on the segmentation results of a template neural network model. The template neural network model is segmented based on a specified segmentation strategy. For example, the specified segmentation strategy can be obtained based on a search and optimization algorithm. Alternatively, the specified segmentation strategy can be obtained based on an artificial expert strategy. Alternatively, the specified segmentation strategy can be obtained based on other methods, which are not specifically limited in the embodiments of the present application. The computing device can obtain a structural template by extracting structural features or node features of the template neural network model. Compared to structural features, node features also include information reflecting the properties of the computing node itself, such as the computing logic of the computing node. The structural template can be all model structures in the template neural network model. Alternatively, the structural template is the model structure after screening the model structures in the template neural network model. The structural template can be obtained by splitting the template neural network model based on structural features. In addition, to ensure the optimization benefits of the segmentation strategy, the segmentation strategies recorded in the knowledge base are all strategies that have been verified to achieve optimization benefits. Here, the template neural network model refers to a neural network model that has been segmented and at least part of the model structure of the neural network model is recorded in the knowledge base. The model structure of the template neural network model recorded in the knowledge base is the structural template. The following is an exemplary description of the process of obtaining the structure template:

[0118] For example, after the computing device obtains the segmented neural network model, it can obtain the segmentation results of each computing node in the neural network model, and group the computing nodes in the neural network model according to the segmentation results to obtain multiple model structures, and then optimize the multiple model structures to obtain one or more structural templates. Optionally, the computing device can group the computing nodes in the neural network model based on a union-find set (disjoint-set data structure). A union-find set is a tree-type data structure used to handle the merging and querying of some disjoint sets. When grouping based on a union-find set, the tree-type data structure can be used to represent the entire neural network model, and then the computing nodes in the same tree can be divided into the same group, and the computing nodes that are no longer in the same tree can be divided into different groups. By filtering with a union-find set, the scale of the computational graph can be reduced.

[0119] In one implementation, after obtaining multiple model structures, the multiple model structures can be optimized based on the principle of segmentation convergence points and / or segmentation consistent areas, and the optimized model structure is a structural template. The principle based on segmentation convergence points means that the structural template needs to meet the following requirements: all computing nodes except the first computing node (also called the first computing node) and the last computing node in the optimized model structure are segmented. In this way, for multiple computing nodes in any optimized model structure, except the first computing node and the last computing node, the remaining computing nodes are all computing nodes with segmentation reference significance, so the model structure can be determined as a structural template for reference when segmenting other model structures. Even the first computing node and the last computing node are also segmented. Accordingly, the optimization process can be regarded as a process of discarding invalid computing nodes that are not segmented in the model structure. For example, Figure 15 is a schematic diagram of a process for obtaining a structural template provided in an embodiment of the present application. The model structure obtained by splitting is shown on the left in Figure 15. In this model structure, the first computing node (i.e., the computing node for executing op type 1), the second computing node (i.e., the computing node for executing op type 2), the last computing node (i.e., the computing node for executing op type 6), and the second-to-last computing node (i.e., the computing node for executing op type 5) are not split. The third computing node in this model structure is split into two computing nodes. The right side of Figure 15 is a schematic diagram of the optimized model structure, that is, the right side of Figure 15 is a schematic diagram of the structural template obtained after the optimization of the model structure. According to Figure 15, it can be seen that this structural template reduces the first and last computing nodes that are not split in the model structure compared to the model structure on the left side of Figure 15.

[0120] The principle of consistent partitioning requires that the optimized model structure meet the following requirements: the number of partitions between different computational nodes in the optimized model structure is an integer multiple. According to the operational logic of neural network models, data is transferred between computationally related computational nodes. When a computational node is partitioned, data originally provided to the node will continue to be provided to the multiple computational nodes derived from it, and the computational nodes that originally received data from the node will continue to receive data from the multiple computational nodes derived from it. Computational nodes within the same model structure are all associated, so the number of partitions between different computational nodes within the same model structure will be an integer multiple. When optimizing the model structure, consider computational nodes that are not continuously partitioned. For any computational node that is not partitioned, check whether its input node is partitioned into a single partition. If not, reassign the input node and the computational node to the same model structure. Similarly, check whether the output node of the computation node is split into 1 part. If the output node is not split into 1 part, divide the output node and the computation node into the same model structure. Similarly, the structural template of the template neural network model can be obtained. For example, in a certain neural network model, the number of parts of the computation node is: [1,1,1,1,1,1,2,2,8,8,8,2,1,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8,8]. According to the principle of dividing the consistent area, model structure 1 composed of computing nodes with split numbers of 2, 2, 8, 8, 8, 2 respectively, and model structure 2 composed of computing nodes with split numbers of 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8 respectively can be extracted, and model structure 1 and model structure 2 can be determined as structural templates.

[0121] Step 1402: Create a knowledge base based on the structure template and its segmentation strategy.

[0122] After obtaining the structural template, you can continue to obtain the segmentation strategy of the calculation nodes in the structural template to obtain the segmentation strategy of the structural template. Then, based on the structural template and its segmentation strategy, create a knowledge base. Optionally, before recording the structural template in the knowledge base, you can optionally perform a serialization operation on the structural template to convert the structural template into a byte sequence. Accordingly, before matching based on the structural template, it is also necessary to perform a deserialization operation on the structural template obtained from the knowledge base to restore the byte sequence to the original representation of the structural template. It should be noted that in the application process of the knowledge base, the segmentation results of the neural network model can also be used to enrich the knowledge base, and its process can refer to the process of creating the knowledge base accordingly.

[0123] Figure 16 is a schematic diagram of a process of segmenting a neural network model provided by an embodiment of the present application. As shown in Figure 16, the method includes two stages.

[0124] Phase 1: The phase of establishing knowledge base.

[0125] After the computing device obtains the model file of the template neural network model, it first converts it into a computational graph. The computational graph is then split based on topological features to obtain multiple model structures. Multiple model structures are then optimized, such as deleting unsplit computing nodes in the model structure to obtain multiple structural templates of the neural network model. The optimization benefit of the splitting strategy of each structural template is verified by the verification system to ensure that each structure is a valid structure that is being optimized. For example, the structural templates before and after splitting are deployed and implemented on the chip respectively, and their respective benefits are obtained. When the benefit of the structural template after splitting is greater than the benefit before splitting, the structural template is determined to be a valid structure that is being optimized and can be added to the knowledge base. Otherwise, the structural template cannot be added to the knowledge base. Then, a knowledge base is created based on the structural template and its splitting strategy.

[0126] The second stage: the application stage of the knowledge base.

[0127] In the process of knowledge base application, after obtaining the structural template from the knowledge base, the structural template is matched with the neural network model to be segmented (i.e., the computational graph). When the target model structure in the neural network model matches the target structure template, the segmentation strategy of the target structure template is obtained from the knowledge base, and then the target model structure is segmented based on the segmentation strategy of the target structure template, and the other model structures in the neural network model are segmented according to the logic, and then the segmented target model structure and the model structure that does not need to be segmented are spliced ​​according to the order of each model structure in the neural network model to obtain the segmented neural network model. Then, an application that can be deployed on the chip is generated based on the segmented model, and then the application is deployed on the chip.

[0128] In one implementation, the neural network model segmentation method provided in the embodiments of the present application can be implemented using multiple functional modules. For example, as shown in FIG17 , a computing device includes a strategy generation module, a model feature extraction module, a knowledge base module, a query matching module, and a model segmentation module. The neural network model segmentation method is implemented through the collaborative operation of these modules.

[0129] The strategy generation module is used to obtain a structural template based on a template neural network model. For example, the strategy generation module determines multiple alternative segmentation strategies of the template neural network model based on a search and tuning algorithm, and then verifies the multiple alternative segmentation strategies, and obtains the optimal segmentation strategy from the multiple alternative segmentation strategies based on the verification results, and provides the optimal segmentation strategy to the model feature extraction module. For example, the strategy for segmenting the template neural network in the above step 1401 is provided by the strategy generation module. Optionally, the neural network model can be a model such as Tensorflow, pytorch, or other neural network models, and the embodiments of the present application do not specifically limit it.

[0130] The model feature extraction module is used to segment the template neural network model based on the optimal segmentation strategy and extract the structural template based on the template neural network model before and after segmentation. For example, the structural template is obtained by extracting the structural features or node features of the template neural network model. For example, the process of segmenting the template neural network in step 1401 and obtaining the structural template based on the segmentation results is performed by the model feature extraction module.

[0131] The knowledge base module is used to create a knowledge base based on the structure template and its segmentation strategy. For example, the above step 1402 can be performed by the knowledge base module.

[0132] The query matching module is used to obtain the optimal segmentation strategy for the neural network model to be segmented based on the neural network model to be segmented and the knowledge base, and provide the segmentation strategy to the model segmentation module. For example, the above steps 501 to 503 can be performed by the query matching module.

[0133] The model segmentation module is used to segment the neural network model to be segmented based on the segmentation strategy provided by the query matching module. When multiple target model structures include model structures with the same computing nodes, the model segmentation module is further used to determine a target model structure combination and segment the target model structure combination based on the segmentation strategy that matches the target model structure combination. For example, the above step 504 can be performed by the model segmentation module.

[0134] In summary, in the segmentation method of the neural network model provided in the embodiment of the present application, since the segmentation strategies recorded in the knowledge base can all obtain optimization benefits, when the target model structure in the neural network model matches the target structure template, the target model structure is segmented based on the segmentation strategy of the target structure template, ensuring that the target model structure after segmentation can obtain optimization benefits compared to before segmentation, so that segmentation of the neural network model can obtain positive benefits, thereby improving the performance of the neural network model. For example, after the present application segments the neural network model, it can reduce the pressure of caching the data of the computing node to the cache without affecting the performance of the computing node in the neural network model. In addition, the segmentation process can be realized automatically, which can reduce the workload and debugging difficulty of segmenting the neural network model.

[0135] It should be noted that the order of the steps in the neural network model segmentation method provided in the embodiments of the present application can be adjusted appropriately, and the number of steps can be increased or decreased accordingly. Any person skilled in the art who can easily conceive of variations within the technical scope disclosed in this application should be included in the scope of protection of this application, and therefore will not be described in detail.

[0136] The following describes the virtual device in the embodiment of the present application by way of example.

[0137] The above introduces the method for segmenting the neural network model of the embodiment of the present application. Corresponding to the above method, the embodiment of the present application also provides a segmentation device for the neural network model. Figure 18 is a structural schematic diagram of a segmentation device for a neural network model provided by an embodiment of the present application. Based on the following multiple components shown in Figure 18, the segmentation device for the neural network model shown in Figure 18 can perform all or part of the operations shown in Figures 5 and 14 above. It should be understood that the device may include more additional components than the components shown or omit some of the components shown therein, and the embodiment of the present application is not limited to this. As shown in Figure 18, the segmentation device 180 of the neural network model may include:

[0138] Matching module 1801 is used to match the neural network model to be segmented with multiple structural templates in the knowledge base to obtain a target structural template that matches the target model structure in the neural network model. The knowledge base also records the segmentation strategy of each structural template. The segmentation strategy indicates the way to segment the structural template. The segmentation strategies recorded in the knowledge base can all achieve optimization benefits. The model structure and structural template both include multiple computing nodes with specified connection relationships.

[0139] The first acquisition module 1802 is used to acquire the segmentation strategy of the target structure template.

[0140] The segmentation module 1803 is used to segment the target model structure based on the segmentation strategy of the target structure template to obtain a segmented neural network model.

[0141] Optionally, matching the target model structure with the target structure template includes: structural features of the target model structure are identical to structural features of the target structure template.

[0142] Optionally, the structural features include one or more of the following: topological features and connection features, where the topological features indicate the execution order between computing nodes, and the connection features indicate the connection relationship between computing nodes.

[0143] Optionally, when multiple target model structures in the neural network model are respectively matched with multiple target structure templates in the knowledge base, the segmentation module 1803 is specifically used to: when there is no model structure including the same computing node among the multiple target model structures, segment the target model structures that match the target structure template based on the segmentation strategies of the multiple target structure templates; when there is a model structure including the same computing node among the multiple target model structures, filter the multiple target model structures so that there is no model structure including the same computing node among the multiple target model structures that have been filtered, and segment the filtered target model structures based on the segmentation strategies of the target structure templates that match the filtered target model structures.

[0144] Optionally, the segmentation module 1803 is specifically used to: combine multiple target model structures to obtain multiple model structure combinations, where the model structure combination includes some target model structures in the multiple target model structures, and the model structure set does not contain repeated computing nodes; based on the optimization benefits of the segmentation strategy of multiple target structure templates, respectively obtain the benefits of the multiple model structure combinations; based on the benefits of the multiple model structure combinations, determine the target model structure combination in the multiple model structure combinations, where the target model structure in the target model structure combination is the screened target model structure.

[0145] Optionally, the matching module 1801 is specifically used to: split the neural network model based on the computing nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model; and match the multiple model structures with the structural template in the knowledge base.

[0146] Optionally, the matching module 1801 is specifically used to: delete the computing nodes that do not conform to the structural features of all the structural templates in the knowledge base from all the computing nodes of the neural network model, and obtain multiple model structures.

[0147] Optionally, the matching module 1801 is specifically configured to: screen a plurality of model structures based on the number of computing nodes included in the structure template in the knowledge base; and match the screened model structures with the structure template in the knowledge base.

[0148] Optionally, as shown in FIG19 , the neural network model segmentation device 180 further includes:

[0149] The second acquisition module 1804 is used to acquire the structure template based on the segmentation result of the template neural network model.

[0150] The creation module 1805 is used to create a knowledge base based on the structure template and its segmentation strategy.

[0151] Optionally, the structural template satisfies one or more of the following: the number of shares of different computing nodes in the structural template is an integer multiple; and all computing nodes except the first computing node and the last computing node in the structural template are split.

[0152] Optionally, both the first compute node and the last compute node are split.

[0153] For the detailed working processes of matching module 1801, first acquisition module 1802, segmentation module 1803, second acquisition module 1804, and creation module 1805, please refer to the description in the previous method embodiment. For example, matching module 1801 matches the neural network model to be segmented with multiple structural templates in the knowledge base in the manner described in step 502 above; first acquisition module 1802 obtains the segmentation strategy of the target structural template in the manner described in step 502 above; segmentation module 1803 obtains the segmented neural network model in the manner described in step 504 above; second acquisition module 1804 obtains the structural template in the manner described in step 1401 above; and creation module 1805 creates the knowledge base in the manner described in step 1402 above. This embodiment of the present application will not be repeated here.

[0154] In summary, in the neural network model segmentation device provided in the embodiment of the present application, the matching module first matches the neural network model to be segmented with multiple structural templates in the knowledge base to obtain a target structural template that matches the target model structure in the neural network model; then, the first acquisition module obtains the segmentation strategy of the target structural template; the segmentation module segments the target model structure based on the segmentation strategy of the target structural template to obtain a segmented neural network model. Among them, the model structure and the structural template both include multiple computing nodes with specified connection relationships. The knowledge base records multiple structural templates and the segmentation strategy of each structural template. The segmentation strategy indicates the way to segment the structural template, and the segmentation strategies recorded in the knowledge base can all achieve optimization benefits. Obtaining optimization benefits here means obtaining positive benefits. In addition, the segmentation process can be automated, which can reduce the workload and debugging difficulty of segmenting the neural network model.

[0155] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the various components described above can refer to the corresponding contents in the aforementioned method embodiments and will not be repeated here.

[0156] The following is an example of the basic hardware structure involved in the embodiments of the present application.

[0157] An embodiment of the present application provides a computing device. The computing device is used to implement some or all of the functions of the segmentation method of the neural network model provided in the embodiment of the present application. Figure 20 is a schematic diagram of the structure of a computing device provided in the embodiment of the present application. As shown in Figure 20, the computing device 2000 includes a processor 2001, a memory 2002, a communication interface 2003 and a bus 2004. Among them, the processor 2001, the memory 2002, and the communication interface 2003 are connected to each other through the bus 2004.

[0158] Processor 2001 may include a general-purpose processor and / or a dedicated hardware chip. A general-purpose processor may include: a central processing unit (CPU), a microprocessor or a graphics processing unit (GPU). The CPU is, for example, a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A dedicated hardware chip is a hardware module for high-performance processing. The dedicated hardware chip includes at least one of a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or a network processor (NP). Processor 2001 may also be an integrated circuit chip with signal processing capabilities. During implementation, some or all of the functions of the segmentation method of the neural network model of the present application may be completed by the hardware integrated logic circuit in the processor 2001 or instructions in the form of software.

[0159] Memory 2002 is used to store computer programs, which include an operating system 2002a and executable code (i.e., program instructions) 2002b. Memory 2002 may be, for example, a read-only memory or other type of static storage device capable of storing static information and instructions, or a random access memory or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory, a read-only optical disc or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired executable code in the form of instructions or data structures and accessible by a computer, but not limited to these. For example, memory 2002 is used to store an outbound port queue, etc. Memory 2002 may exist independently and be connected to processor 2001 via bus 2004. Alternatively, memory 2002 and processor 2001 may be integrated together. The memory 2002 can store executable code. When the executable code stored in the memory 2002 is executed by the processor 2001, the processor 2001 is used to perform part or all of the functions of the neural network model segmentation method provided in the embodiment of the present application. For the implementation of the processor 2001 executing this process, please refer to the relevant description in the aforementioned embodiment. The memory 2002 may also include software modules and data required for other running processes such as the operating system.

[0160] The communication interface 2003 uses a transceiver module, such as, but not limited to, a transceiver, to communicate with other devices or communication networks. For example, the communication interface 2003 can be any one or a combination of the following devices: a network interface (such as an Ethernet interface), a wireless network card, or other device with network access capabilities.

[0161] Bus 2004 is any type of communication bus used to interconnect the internal components of a computing device (e.g., memory 2002, processor 2001, and communication interface 2003). For example, a system bus is provided. The embodiments of this application illustrate the example of interconnecting the aforementioned components within a computing device via bus 2004. Alternatively, the aforementioned components within computing device 2000 may also be communicatively connected to each other using other connection methods besides bus 2004. For example, the aforementioned components within computing device 2000 may be interconnected via an internal logical interface.

[0162] It should be noted that the above-mentioned multiple devices can be respectively arranged on independent chips, or at least partially or completely arranged on the same chip. Whether each device is independently arranged on different chips or integrated on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation form of the above-mentioned devices. The descriptions of the processes corresponding to the above-mentioned figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.

[0163] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product that provides a program development platform includes one or more computer instructions. When these computer program instructions are loaded and executed on a computing device, the functions of the segmentation method of the neural network model provided in the embodiments of the present application are fully or partially implemented.

[0164] Furthermore, computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium stores computer program instructions that provide a program development platform.

[0165] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0166] Optionally, the structure of at least one computing device included in the computing device cluster can refer to the computing device 2000 shown in Figure 20. The memory 2002 in one or more computing devices 2000 in the computing device cluster can store the same instructions for executing the segmentation method of the neural network model.

[0167] In some possible implementations, the memory 2002 of one or more computing devices 2000 in the computing device cluster may also store partial instructions for executing the neural network model segmentation method. In other words, the combination of one or more computing devices 2000 can jointly execute the instructions for executing the neural network model segmentation method.

[0168] It should be noted that the memory 2002 in different computing devices 2000 in the computing device cluster can store different instructions, each for executing a portion of the functions of the neural network model segmentation device. In other words, the instructions stored in the memory 2002 in different computing devices 2000 can implement the functions of one or more of the matching module 1801, the first acquisition module 1802, the segmentation module 1803, the second acquisition module 1804, and the creation module 1805.

[0169] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 21 illustrates a possible implementation. As shown in FIG. 21 , two computing devices 2100A and 2100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, computing devices 2100A and 2100B include a bus 2102, a processor 2104, a memory 2106, and a communication interface 2108. The memory 2106 in computing device 2100A stores instructions for executing the functions of matching module 1801, first acquisition module 1802, and segmentation module 1803. Simultaneously, the memory 2106 in computing device 2100B stores instructions for executing the functions of second acquisition module 1804 and creation module 1805.

[0170] It should be understood that the functions of the computing device 2100A shown in FIG21 may also be performed by multiple computing devices 2100. Similarly, the functions of the computing device 2100B may also be performed by multiple computing devices 2100. Furthermore, the deployment method of the modules for implementing the segmentation method of the neural network model in the computing devices may also be adjusted according to application requirements.

[0171] An embodiment of the present application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes program instructions. When the program instructions are executed on a computing device, the computing device implements the neural network model segmentation method provided in the embodiment of the present application.

[0172] An embodiment of the present application also provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer implements the neural network model segmentation method provided by the embodiment of the present application.

[0173] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0174] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.

[0175] In the embodiments of the present application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "plurality" refers to two or more, unless otherwise expressly limited.

[0176] In this application, the term "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0177] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the concepts and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A segmentation method for a neural network model, characterized in that: The method comprises: Matching the neural network model to be segmented with a plurality of structural templates in a knowledge base to obtain a target structural template that matches the target model structure in the neural network model, wherein the knowledge base also records segmentation strategies for each structural template, wherein the segmentation strategies indicate a method for segmenting the structural template, and the segmentation strategies recorded in the knowledge base can all achieve optimization benefits, and the model structure and the structural template both include a plurality of computing nodes with a specified connection relationship; Obtaining a segmentation strategy for the target structure template; Based on the segmentation strategy of the target structure template, the target model structure is segmented to obtain a segmented neural network model.

2. The method according to claim 1, characterized in that The matching of the target model structure with the target structure template includes: a structural feature of the target model structure is the same as a structural feature of the target structure template.

3. The method according to claim 2, characterized in that The structural features include one or more of the following: topological features and connection features, wherein the topological features indicate the execution order between computing nodes, and the connection features indicate the connection relationship between computing nodes.

4. The method according to any one of claims 1 to 3, characterized in that: When multiple target model structures in the neural network model are matched with multiple target structure templates in the knowledge base respectively, the segmentation strategy based on the target structure template segments the target model structure, including: When there is no model structure including the same computing node among the multiple target model structures, segmenting the target model structures matching the target structure template based on the segmentation strategies of the multiple target structure templates respectively; When there are model structures including the same computing nodes among the multiple target model structures, the multiple target model structures are screened so that there are no model structures including the same computing nodes among the screened multiple target model structures, and the screened target model structures are segmented based on the segmentation strategies of the target structure templates that match the screened target model structures.

5. The method according to claim 4, characterized in that The screening of the multiple target model structures so that no model structure including the same computing node exists among the screened multiple target model structures comprises: Combining the multiple target model structures to obtain multiple model structure combinations, wherein the model structure combinations include some of the target model structures in the multiple target model structures, and the model structure set does not contain repeated computing nodes; Based on the optimization benefits of the segmentation strategies of the multiple target structure templates, respectively obtaining benefits of the multiple model structure combinations; Based on the benefits of the multiple model structure combinations, a target model structure combination is determined from the multiple model structure combinations, wherein the target model structure in the target model structure combination is a screened target model structure.

6. The method according to any one of claims 1 to 5, characterized in that: Before matching the neural network model to be segmented with a plurality of structural templates in the knowledge base, the method further includes: Splitting the neural network model based on the computing nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model; The matching of the neural network model to be segmented with multiple structural templates in the knowledge base includes: The multiple model structures are matched with structure templates in the knowledge base.

7. The method according to claim 6, characterized in that The neural network model is split based on the computing nodes included in the structure template in the knowledge base to obtain multiple model structures of the neural network model, including: Among all the computing nodes of the neural network model, the computing nodes that do not conform to the structural features of all the structural templates in the knowledge base are deleted to obtain the multiple model structures.

8. The method according to claim 6 or 7, characterized in that Before matching the multiple model structures with the structure templates in the knowledge base, the method further includes: Screening the multiple model structures based on the number of computing nodes included in the structure template in the knowledge base; The matching of the multiple model structures with the structure templates in the knowledge base includes: The screened model structures are matched with the structural templates in the knowledge base.

9. The method according to any one of claims 1 to 8, characterized in that: Before matching the neural network model to be segmented with a plurality of structural templates in the knowledge base, the method further includes: Based on the segmentation result of the template neural network model, obtaining the structure template; The knowledge base is created based on the structure template and its segmentation strategy.

10. The method according to claim 9, characterized in that The structural template satisfies one or more of the following: The number of divisions of different computing nodes in the structure template is in integer multiples; Furthermore, all computing nodes except the first computing node and the last computing node in the structure template are split.

11. The method according to claim 10, characterized in that The first computing node and the last computing node are both split.

12. A neural network model segmentation device, characterized in that: The device comprises: A matching module is used to match the neural network model to be segmented with a plurality of structural templates in a knowledge base, and obtain a target structural template that matches the target model structure in the neural network model, wherein the knowledge base also records segmentation strategies of each structural template, wherein the segmentation strategies indicate a method for segmenting the structural template, and the segmentation strategies recorded in the knowledge base can all achieve optimization benefits, and the model structure and the structural template both include a plurality of computing nodes with a specified connection relationship; A first acquisition module, used to acquire a segmentation strategy of the target structure template; A segmentation module is used to segment the target model structure based on the segmentation strategy of the target structure template to obtain a segmented neural network model.

13. The device according to claim 12, characterized in that The matching of the target model structure with the target structure template includes: a structural feature of the target model structure is the same as a structural feature of the target structure template.

14. The device according to claim 13, characterized in that The structural features include one or more of the following: topological features and connection features, wherein the topological features indicate the execution order between computing nodes, and the connection features indicate the connection relationship between computing nodes.

15. The device according to any one of claims 12 to 14, characterized in that: When the multiple target model structures in the neural network model are matched with the multiple target structure templates in the knowledge base respectively, the segmentation module is specifically used to: When there is no model structure including the same computing node among the multiple target model structures, segmenting the target model structures matching the target structure template based on the segmentation strategies of the multiple target structure templates respectively; When there are model structures including the same computing nodes among the multiple target model structures, the multiple target model structures are screened so that there are no model structures including the same computing nodes among the screened multiple target model structures, and the screened target model structures are segmented based on the segmentation strategies of the target structure templates that match the screened target model structures.

16. The device according to claim 15, characterized in that The segmentation module is specifically used for: Combining the multiple target model structures to obtain multiple model structure combinations, wherein the model structure combinations include some of the target model structures in the multiple target model structures, and the model structure set does not contain repeated computing nodes; Based on the optimization benefits of the segmentation strategies of the multiple target structure templates, respectively obtaining benefits of the multiple model structure combinations; Based on the benefits of the multiple model structure combinations, a target model structure combination is determined from the multiple model structure combinations, wherein the target model structure in the target model structure combination is a screened target model structure.

17. The device according to any one of claims 12 to 16, characterized in that: The matching module is specifically used for: Splitting the neural network model based on the computing nodes included in the structural template in the knowledge base to obtain multiple model structures of the neural network model; The multiple model structures are matched with structure templates in the knowledge base.

18. The device according to claim 17, characterized in that The matching module is specifically used for: Among all the computing nodes of the neural network model, the computing nodes that do not conform to the structural features of all the structural templates in the knowledge base are deleted to obtain the multiple model structures.

19. The device according to claim 17 or 18, characterized in that The matching module is specifically used for: Screening the multiple model structures based on the number of computing nodes included in the structure template in the knowledge base; The screened model structures are matched with the structural templates in the knowledge base.

20. The device according to any one of claims 12 to 19, characterized in that The device also includes: A second acquisition module is used to acquire the structure template based on the segmentation result of the template neural network model; A creation module is used to create the knowledge base based on the structure template and its segmentation strategy.

21. The device according to claim 20, characterized in that The structural template satisfies one or more of the following: The number of divisions of different computing nodes in the structure template is in integer multiples; Furthermore, all computing nodes except the first computing node and the last computing node in the structure template are split.

22. The device according to claim 21, characterized in that The first computing node and the last computing node are both split.

23. A computing device cluster, characterized in that: The method comprises at least one computing device, each computing device comprises a processor and a memory, and the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 11.

24. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 11.

25. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Convolutional neural network calculation method and system

    CN110633785A

  • Dynamic load balancing method for model parallelism of neural network

    CN114217944A

  • Model segmentation method and related equipment thereof

    CN114707643A

  • Neural network compiling optimization method and related apparatus

    WO2022087788A1

  • Data processing method and apparatus

    WO2023122854A1