Memory management method and device, equipment and storage medium
By determining the initial computing node of each execution step according to the target execution order in the target neural network model, and combining the memory footprint of the jump-to-connected starting node and the initial computing node, the accuracy of memory footprint calculation during the model optimization process is solved, and the precise calculation of the memory footprint of each execution step is achieved.
Patent Information
- Application Number
- CN202510206502.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
During the model optimization process, it is difficult to accurately calculate the memory occupied by the target neural network model at each execution step, especially in the presence of a jump connection.
By determining the initial computing node of each execution step according to the order of target execution, and when there is a hopping connection, the total memory footprint of the target neural network model is calculated.
Accurate calculation of the memory usage of the target neural network model in each execution step is achieved, especially in the presence of a jump connection, which improves the accuracy of memory management.
Smart Images

Figure CN119988032A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to model optimization technology, and relate to but are not limited to a memory management method and apparatus, equipment, and storage medium. Background Art
[0002] In the process of optimizing and improving the model, it is usually necessary to obtain the memory occupied by each execution step in the model during the running process. Therefore, a solution that can accurately calculate the memory occupied by the model at each execution step is urgently needed. Summary of the invention
[0003] In view of this, the memory management method, device, equipment, and storage medium provided in the embodiments of the present application can calculate the memory occupied by the target neural network model in each execution step. The memory management method, device, equipment, and storage medium provided in the embodiments of the present application are implemented as follows:
[0004] In one aspect of an embodiment of the present application, a memory management method is provided, which is applied to an electronic device, wherein the electronic device runs a target neural network model, the target neural network model includes multiple computing nodes, and the multiple computing nodes run in a target execution order to implement the function of the target neural network model, the method comprising:
[0005] According to the target execution order, determine the initial computing nodes included in each execution step;
[0006] In the case where a target neural network model has a skip connection, the memory occupied by the target neural network model is determined based on the memory occupied by the starting node of the skip connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, wherein the execution steps of the starting node and the ending node of the skip connection are discontinuous, and the computing nodes running in the target execution step of the target neural network model include: the starting node of the skip connection and the initial computing node of the target execution step; and the target execution step is an execution step located between the starting execution step corresponding to the starting node and the ending execution step corresponding to the ending node.
[0007] In another aspect of the embodiment of the present application, a memory management device is provided, which is applied to an electronic device, wherein the electronic device runs a target neural network model, wherein the target neural network model includes a plurality of computing nodes, and the plurality of computing nodes run in a target execution order to implement the function of the target neural network model, wherein the device includes: a node determination module, and a memory calculation module;
[0008] A node determination module, used to determine the initial computing nodes included in each execution step according to the target execution order;
[0009] A memory computing module is used to determine the memory occupied by the target neural network model when a jump connection exists in the target neural network model, based on the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, wherein the execution steps of the starting node and the ending node of the jump connection are discontinuous, and the computing nodes running in the target execution step of the target neural network model include: the starting node of the jump connection and the initial computing node of the target execution step; the target execution step is an execution step located between the starting execution step corresponding to the starting node and the ending execution step corresponding to the ending node.
[0010] The computer device provided in the embodiment of the present application includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the program, the method of the embodiment of the present application is implemented.
[0011] The computer-readable storage medium provided in the embodiment of the present application has a computer program stored thereon, and when the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0012] In the memory management method, apparatus, device, and storage medium provided in the embodiments of the present application, the initial computing node included in each execution step can be determined according to the target execution order; in the case where the target neural network model has a jump connection, the memory occupied by the target neural network model can be determined according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime. Among them, the target execution step can be an execution step located between the starting execution step corresponding to the starting node and the terminating execution step corresponding to the terminating node. In the process of calculating the occupied memory, the calculation can be performed according to the computing node running at the target execution step of the target neural network model, that is, the calculation of the occupied memory corresponding to the target execution step can be implemented according to the starting node of the jump connection and the initial computing node of the target execution step, so that the memory occupied by the target neural network model in each execution step can be accurately determined. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0014] Figure 1 A schematic diagram of the structure of a neural network model provided in an embodiment of the present application;
[0015] Figure 2A flowchart of a memory management method provided in an embodiment of the present application;
[0016] Figure 3 A schematic diagram of a neural network structure with a skip connection provided in an embodiment of the present application;
[0017] Figure 4 A schematic diagram of a neural network structure with multiple skip connections provided in an embodiment of the present application;
[0018] Figure 5 A schematic diagram of a process for determining the memory occupied by a neural network model provided in an embodiment of the present application;
[0019] Figure 6 A schematic diagram of a determination process for determining whether a neural network model has a skip connection provided in an embodiment of the present application;
[0020] Figure 7 A schematic diagram of a node pair provided in an embodiment of the present application;
[0021] Figure 8 A schematic diagram of a process for determining target execution steps provided in an embodiment of the present application;
[0022] Fig. 9 A schematic diagram of the screening results of node pairs provided in an embodiment of the present application;
[0023] Fig.10 A schematic diagram of the result of expanding the neural network model provided in the embodiments of the present application;
[0024] Fig.11 Another schematic diagram of the memory management method provided in the embodiment of the present application;
[0025] Fig.12 It is another flowchart of the memory management method provided in the embodiment of the present application;
[0026] Fig.13 A schematic diagram of the calculation results of the memory usage of each step provided in the embodiments of the present application;
[0027] Fig.14 A schematic diagram of the structure of a memory management device provided in an embodiment of the present application;
[0028] Fig.15 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0031] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0033] It should be noted that the present application may provide an electronic device, wherein the electronic device may include but is not limited to mobile phones, wearable devices (such as smart watches, smart bracelets, smart glasses, etc.), tablet computers, laptop computers, vehicle terminals, PCs (Personal Computers), etc. The functions implemented by the method can be implemented by calling program codes by a processor in the electronic device, and of course the program codes can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.
[0034] A target neural network model can be run on the electronic device. The target neural network model includes multiple computing nodes. The multiple computing nodes run in a target execution order to implement the functions of the target neural network model.
[0035] Among them, the target neural network model can be, for example, a deep neural network model, a convolutional neural network model and other types of neural network models. There is no restriction on the specific type here, and the corresponding neural network model can be adapted according to actual needs.
[0036] A computing node may refer to an operator (node) used for calculation in the neural network model. For example, the neural network model may be divided into multiple levels, each neural network level may include multiple computing nodes, and these computing nodes may have a corresponding target execution order. For example, node 1 may be executed first, then node 2, and finally node 3.
[0037] The target neural network model provided in the embodiments of the present application is explained below through the structure of a specific neural network model.
[0038] Figure 1 This is a schematic diagram of the structure of the neural network model provided in the embodiment of the present application, please refer to Figure 1 , Figure 1 The neural network model shown in may include 7 computing nodes. During the actual operation, these computing nodes may be executed sequentially or in parallel.
[0039] by Figure 1 Taking the execution order shown in as an example, from node 1 to node 2, node 5 to node 7, they are all calculation stages executed in sequence. For this type of computing nodes, in the process of determining the memory occupied by the neural network model, the memory occupied by each computing node can be used as the memory occupied by the neural network model at that moment.
[0040] However, in the actual implementation process, there may be a situation where a skip connection is executed between node 2 and node 5. For such computing nodes, it is not accurate to use the memory occupied by the neural network model at the corresponding moment only based on the memory occupied by each computing node. Therefore, there is an urgent need for a solution that can accurately calculate the memory occupied by the model at each execution step, especially for the case where there are skip connection steps.
[0041] In order to solve the above problems existing in the related art, a memory management method is provided in an embodiment of the present application, through which the memory occupied by the model in each execution step can be accurately calculated.
[0042] The following is an explanation of one feasible implementation process of the memory management method provided in the embodiment of the present application.
[0043] Figure 2 Please refer to the flowchart of the memory management method provided in the embodiment of the present application. Figure 2 , the method comprising:
[0044] S210: Determine the initial computing nodes included in each execution step according to the target execution order.
[0045] It should be noted that the execution subject of the method may be the above-mentioned electronic device.
[0046] Among them, the target execution order refers to the execution order of each computing node in the target neural network model, and the execution step refers to each corresponding step in the target execution order, among which different computing nodes may have the same execution steps.
[0047] Please refer to the above Figure 1 , Figure 1 Node 3 and node 4 are parallel nodes, and both can be used as initial computing nodes for executing step 3.
[0048] by Figure 1 Taking the target neural network model of 7 computing nodes shown as an example, the initial computing node included in executing step 1 is node 1, the initial computing node included in executing step 2 is node 2, the initial computing nodes included in executing step 3 are node 3 and node 4, the initial computing node included in executing step 4 is node 5, the initial computing node included in executing step 5 is node 6, and the initial computing node included in executing step 6 is node 7.
[0049] That is to say, the execution step to which each computing node belongs can be determined according to the target execution order composed of each computing node in the target neural network, thereby determining the initial computing node included in each execution step.
[0050] S220: In the case where the target neural network model has a skip connection, the memory occupied by the target neural network model is determined according to the memory occupied by the starting node of the skip connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime.
[0051] It should be noted that a skip connection refers to a connection relationship in which the execution steps of the starting node and the ending node are discontinuous, for example: Figure 1 Node 2 and node 5 are connected, node 2 belongs to execution step 2, node 5 belongs to execution step 4, and execution step 2 and execution step 4 are not continuous. In this case, it can be determined that the connection relationship is a skip connection relationship.
[0052] Optionally, it is possible to determine whether the target neural network model has a skip connection. If there is no skip connection, the memory occupied by the target neural network model can be determined based on the initial computing node included in each execution step. If there is a skip connection, the memory occupied by the target neural network model can be determined based on the memory occupied by the starting node of the skip connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime.
[0053] It should be noted that the target neural network may include: target execution steps and other execution steps, wherein the target execution step is an execution step between the start execution step corresponding to the start node and the end execution step corresponding to the end node. Other execution steps are the remaining execution steps except the target execution step.
[0054] Example: Figure 1 For example, the target execution step is the execution step between execution step 2 corresponding to node 2 and execution step 4 corresponding to node 5, that is, execution step 3.
[0055] It should be noted that in the process of calculating the memory occupied by the target neural network model, different methods can be used for calculation according to different execution steps.
[0056] For example: for other execution steps, the memory occupied by the execution step can be determined based on the memory occupied by the initial computing node; for the target execution step, the memory occupied by the execution step can be determined based on the memory occupied by the initial computing node and the memory occupied by the starting node of the jump connection.
[0057] In one embodiment, the computing nodes of the target neural network model running in the target execution step include: the starting node of the jump connection and the initial computing node of the target execution step.
[0058] It should be noted that since the computing nodes running the target neural network model are not the same in different running stages, the memory occupied by the target neural network model is a variable quantity. The memory occupied by the target neural network model can be divided into the memory occupied by the target neural network model when executing each execution step.
[0059] Among them, each execution step is executed in sequence, for example: first execute execution step 1 of the target neural network model, then execute execution step 2 of the target neural network model, and so on, until executing execution step N of the target neural network model, N is the last execution step of the target neural network model.
[0060] In the memory management method provided in the embodiment of the present application, the initial computing node included in each execution step can be determined according to the target execution order; in the case where the target neural network model has a jump connection, the memory occupied by the target neural network model can be determined according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime. Among them, the target execution step can be an execution step located between the starting execution step corresponding to the starting node and the terminating execution step corresponding to the terminating node. In the process of calculating the occupied memory, the calculation can be performed according to the computing node running at the target execution step of the target neural network model, that is, the calculation of the occupied memory corresponding to the target execution step can be realized according to the starting node of the jump connection and the initial computing node of the target execution step, so that the memory occupied by the target neural network model in each execution step can be accurately determined.
[0061] It should be noted that for different neural network models, the number of skip connections included may be one or more, and no specific limitation is made here. Whether there is a corresponding skip connection can be determined based on the execution relationship of the actual computing nodes. The following explains the impact of different numbers of skip connections on the memory management method using specific neural network structures with one skip connection and with multiple skip connections.
[0062] Figure 3 This is a schematic diagram of a neural network structure with a skip connection provided in an embodiment of the present application. Please refer to Figure 3 , Figure 3 The neural network structure shown in Figure 1 The partial structure composed of nodes 2 to 5 is a neural network structure with a skip connection.
[0063] Figure 3 It includes: node conv_1, node conv_2, node conv_3 and node add_1. During the execution of these four nodes, node conv_1 can be run first, and then nodes conv_2 and conv_3 can be run respectively according to the output of node conv_1. Finally, node add_1 can be run according to the output of node conv_1, node conv_2 and node conv_3.
[0064] Figure 3 The structure of the neural network model shown is a structure with a skip connection, wherein the starting node of the skip connection is the node conv_1, and the ending node of the skip connection is the node add_1.
[0065] for Figure 3The target neural network model shown in the figure includes executing step 1, executing step 2 and executing step 3, wherein the initial computing node included in executing step 1 is node conv_1; the initial computing node included in executing step 2 is node conv_2 and node conv_3; the initial computing node included in executing step 3 is node add_1.
[0066] Among them, executing step 2 is the above-mentioned target execution step.
[0067] Figure 4 This is a schematic diagram of a neural network structure with multiple skip connections provided in an embodiment of the present application. Please refer to Figure 4 , Figure 4 The neural network shown in can be a structure with multiple skip connections.
[0068] Figure 4 It includes: node conv_4, node conv_5, node add_2, node conv_6 and node add_3. During the execution of these five nodes, node conv_1 can be run first, and then nodes conv_2 and conv_3 can be run respectively according to the output of node conv_1, and finally node add can be run according to the output of node conv_1, node conv_2 and node conv_3.
[0069] Figure 4 The structure of the neural network model shown is a structure with multiple skip connections (taking two skip connections as an example), wherein the starting node of the first skip connection is node conv_4, and the ending node of the first skip connection is node add_2; the starting node of the second skip connection is node conv_5, and the ending node of the second skip connection is node add_3.
[0070] for Figure 4 The target neural network model shown in the figure includes executing step 1, executing step 2, executing step 3, executing step 4 and executing step 5, wherein the initial computing node included in executing step 1 is node conv_4; the initial computing node included in executing step 2 is node conv_5; the initial computing node included in executing step 3 is node add_2; the initial computing node included in executing step 4 is node conv_6; the initial computing node included in executing step 5 is node add_3.
[0071] Among them, executing step 2, executing step 3, and executing step 4 are the above-mentioned target execution steps, executing step 2 is a target execution step based on the first jump connection, and executing step 3 and executing step 4 are target execution steps based on the second jump connection.
[0072] Figure 4 Two skip connections are used as an example. In practical applications, the neural network model may have more skip connections.
[0073] A feasible implementation process for determining the memory occupied by a neural network model provided in an embodiment of the present application is explained below.
[0074] Figure 5 Please refer to the flowchart of determining the memory occupied by the neural network model provided in the embodiment of the present application. Figure 5 , according to the memory occupied by the starting node of the skip connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, the memory occupied by the target neural network model is determined, including:
[0075] S510: Determine the computing node on which the target network model runs when executing the target execution step according to the starting node of the jump connection and the initial computing node included in the target execution step.
[0076] It should be noted that the computing node running the target execution step may include the starting node of the jump connection and the initial computing node included in the target execution step.
[0077] Among them, Figure 3 Taking the target neural network model with a jump connection as shown as an example, the target execution step is to execute step 2, wherein the starting node of the jump connection is node conv_1, and the initial computing nodes included in the target execution step are node conv_2 and node conv_3. It can be obtained that the computing nodes running in the target network model when executing the target execution step are node conv_1, node conv_2 and node conv_3.
[0078] by Figure 4 Taking the target neural network model of multiple jump connections shown as an example, the target execution steps are to execute step 2, execute step 3, and execute step 4, wherein the starting node of the jump connection of executing step 2 is node conv_4, and the initial computing node included in executing step 2 is node conv_5. It can be obtained that the computing nodes running on the target network model in executing the target execution step are node conv_4 and node conv_5. The starting node of the jump connection of executing steps 3 and 4 is node conv_5, the initial computing node included in executing step 3 is node add_1, and the initial computing node included in executing step 4 is node conv_6. It can be obtained that the computing nodes running on the target network model in executing step 3 are node conv_4 and node add_1; the computing nodes running on the target network model in executing step 4 are node conv_4 and node conv_5.
[0079] In the above manner, the corresponding running computing nodes under the target execution step of each jump connection can be calculated respectively.
[0080] S520: Determine the memory occupied by the target neural network model at each execution step according to the memory occupied by the computing nodes running the target neural network model at each execution step.
[0081] Each execution step includes a target execution step.
[0082] Optionally, the target neural network model may include a target execution step and other execution steps. The computing nodes running in the target execution step of the target neural network model may be determined according to the method adopted in S510, and the computing nodes running in other execution steps of the target neural network model may be the initial computing nodes included in the execution step.
[0083] After determining the computing node running at each execution step, the memory occupied by the target neural network model at each execution step can be determined based on the memory occupied by the computing node.
[0084] For example, the specific formula for the memory occupied by the target execution step is as follows:
[0085] Si=Si0+St;
[0086] Among them, i is the i-th execution step, Si is the memory occupied by the i-th target execution step, Si0 is the memory occupied by the initial computing node included in the i-th target execution step, and St is the memory occupied by the starting node of the jump connection.
[0087] For example, the specific formula for memory occupied by other execution steps is as follows:
[0088] Sj = Sj0;
[0089] Among them, j is the j-th execution step, Sj is the memory occupied by the j-th other execution step, and Sj0 is the memory occupied by the initial computing node included in the j-th other execution step.
[0090] The above method can be used to determine the memory occupied by the target neural network model at each execution step.
[0091] In the memory management method provided in the embodiment of the present application, the computing node where the target network model runs in the target execution step can be determined according to the starting node of the jump connection and the initial computing node included in the target execution step; the memory occupied by the computing node where the target neural network model runs in each execution step can be determined. Among them, the computing node where the target execution step runs can be determined by the starting node of the jump connection and the initial node of the target execution step, and then the occupied memory can be determined based on the computing node where the target execution step runs; further, the memory occupied by the computing node running in each execution step can be determined, which can ensure the comprehensiveness of the occupied memory and the accuracy of the memory occupied by each execution step.
[0092] The following is an explanation of one feasible implementation method of performing jump connection determination provided in the embodiments of the present application.
[0093] Figure 6 This is a schematic diagram of the determination process of whether a neural network model has a skip connection provided in the embodiment of the present application. Please refer to Figure 6 Before determining the memory occupied by the target neural network model according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, the method further includes:
[0094] S610: Determine the next computing node corresponding to each computing node in the target neural network model in the target execution order.
[0095] It should be noted that after determining multiple computing nodes in the target neural network model, the next computing node corresponding to each computing node can be determined.
[0096] Among them, the node connected to the output end of each computing node according to the target execution order is the next computing node of the computing node. Each computing node can have one or more next computing nodes. If it is a computing node in normal order, the execution steps of the next computing node and the execution steps of the computing node are continuous. If it is a jump-connected computing node, the execution steps of the next computing node and the execution steps of the computing node are discontinuous.
[0097] S620: When the target computing node corresponds to at least two next computing nodes and the execution steps corresponding to the at least two next computing nodes are different, determine that a jump connection exists in the target neural network model.
[0098] The target computing node is the starting node of the jump connection.
[0099] It should be noted that for computing nodes in normal order, the execution steps of the corresponding multiple next computing nodes are the same; for computing nodes with jump connections, the execution steps of the corresponding multiple next computing nodes are different. Whether there is a jump connection in the target neural network model can be determined based on the continuity of the execution steps corresponding to the computing node and the next computing node.
[0100] For example: if a computing node corresponds to executing step 2, and the two next computing nodes correspond to executing step 3, it can be determined that the computing node is not the starting node of the skip connection; if a computing node corresponds to executing step 2, and the two next computing nodes correspond to executing step 3 and step 4 respectively, it can be determined that the computing node is the starting node of the skip connection, and it can be determined that the target neural network model has a skip connection.
[0101] In other words, the relationship between multiple nodes can be used to determine whether the target neural network model has skip connections.
[0102] In the memory management method provided in the embodiment of the present application, the next computing node corresponding to each computing node in the target neural network model in the target execution order can be determined; when the target computing node corresponds to at least two next computing nodes, and the execution steps corresponding to at least two next computing nodes are different, it is determined that there is a jump connection in the target neural network model. Among them, it is possible to determine whether the computing node is the starting node of the jump connection through the relationship between each computing node and the next computing node of the computing node, thereby accurately determining whether there is a jump connection in the target neural network model.
[0103] In one embodiment, determining the next computing node corresponding to each computing node in the target neural network model in the target execution order includes: obtaining all node pairs in the target neural network model.
[0104] The node pair is used to represent any two computing nodes having a connection relationship in the target neural network model, and each node pair includes: a starting node and an ending node.
[0105] It should be noted that each set of node pairs can be extracted from the target neural network model. Figure 1 Taking the neural network model with 7 computing nodes as an example, multiple groups of node pairs can be obtained. Multiple node pairs can be Figure 7 to achieve the displayed content.
[0106] Figure 7 For a schematic diagram of a node pair provided in an embodiment of the present application, please refer to Figure 7 , Figure 7 The multiple node pairs in Figure 1 The neural network model shown is used to implement node pair extraction.
[0107] A node pair can consist of two adjacent nodes, for example: node 1-node 2, node 2-node 3, node 2-node 4, node 2-node 5, node 3-node 5, node 4-node 5, node 5-node 6, node 6-node 7.
[0108] The above multiple node pairs are Figure 1 All the node pairs included in the neural network model shown, Figure 7 As an example, in actual application, the number of nodes in the neural network model may be greater, and there may also be more node pairs.
[0109] The node on the left side of the node pair is the starting node, and the node on the right side is the ending node. A node pair is used as an example for explanation: in node 2-node 3, node 2 is the starting node, and node 3 is the ending node.
[0110] It should be noted that, according to Figure 7 The multiple node pairs obtained as shown in are used to determine whether there is a skip connection in the neural network model.
[0111] In one embodiment, when a target computing node corresponds to at least two next computing nodes, and the execution steps corresponding to the at least two next computing nodes are different, it is determined that a jump connection exists in the target neural network model, including: in each node pair, the same starting node corresponds to at least two different terminating nodes, and the execution steps of at least two different terminating nodes are different, it is determined that a jump connection exists in the target neural network model.
[0112] Among them, for the same starting node, if it corresponds to two different ending nodes, and the execution steps corresponding to the two different ending nodes are different, it can be determined that the starting node is the starting node of the jump connection, and it can be determined that there is a jump connection in the target neural network model.
[0113] For example, Figure 7 Taking the multiple end nodes corresponding to node 2 as an example, the node pairs with node 2 as the starting node include: node 2-node 3, node 2-node 4, node 2-node 5, among which the execution steps corresponding to node 3 and node 4 are the same, both are executing step 3, and the execution step corresponding to node 5 is executing step 4.
[0114] Since the execution steps corresponding to nodes 3 and 5, or nodes 4 and 5, are not the same, it can be determined that node 2 is the starting node of the skip connection, and the skip connection exists in the target neural network model.
[0115] In the memory management method provided in the embodiment of the present application, all node pairs in the target neural network model can be obtained. In each node pair, the same starting node corresponds to at least two different termination nodes, and the execution steps of at least two different termination nodes are different, and it is determined that there is a jump connection in the target neural network model. By obtaining node pairs, it is possible to more quickly and accurately determine whether there is a jump connection in the target neural network model.
[0116] In the actual implementation process, the target execution step can also be determined from the multiple execution steps of the neural network model. The specific determination process of achieving the target execution step is as follows:
[0117] Figure 8 Please refer to the flowchart of the target execution steps provided in the embodiment of this application. Figure 8 , before determining the target network model based on the starting node of the jump connection and the initial computing node included in the target execution step and the computing node running in the target execution step, the method also includes:
[0118] S810: Filter node pairs from all node pairs of the target neural network model to obtain the maximum termination node corresponding to the same starting node.
[0119] Among them, the maximum termination node is the termination node with the latest execution steps in the target neural network model.
[0120] It should be noted that Figure 7 Filter multiple node pairs provided in Figure 7 A part of the node pairs is filtered out from all the node pairs shown in , for example, the unique end node corresponding to the same start node can be filtered out.
[0121] Example: Figure 7 The node pairs with node 2 as the starting node include: node 2-node 3, node 2-node 4 and node 2-node 5, from which the largest termination node can be screened out. Among them, the execution step corresponding to node 3 and node 4 is execution step 3, and the execution step corresponding to node 5 is execution step 4. The execution step corresponding to node 5 is later. Node 5 is the largest termination node corresponding to node 2. The node pair of node 2-node 5 can be screened out. For a node pair with only one termination node, this node pair can be screened out.
[0122] The following is an explanation of the multiple node pairs included after screening provided in the embodiments of the present application.
[0123] Fig. 9 This is a schematic diagram of the screening results of the node pairs provided in the embodiment of the present application, please refer to Fig. 9After filtering the node pairs, we can get the unique terminal node corresponding to each starting node. Figure 7 Taking the multiple nodes shown in the figure as an example, we can get the following Fig. 9 Multiple node pairs are shown.
[0124] For example: it may include node 1-node 2, node 2-node 5, node 3-node 5, node 4-node 5, node 5-node 6, node 6-node 7.
[0125] After obtaining the filtered node pairs, the target execution steps can be determined.
[0126] S820: Determine a target execution step according to the execution step of the start node and the execution step of the maximum end node.
[0127] It should be noted that after obtaining the starting node and the maximum ending node of each node pair through the above screening steps, the target execution step can be determined, and the specific method is as follows:
[0128] If the execution step of the start node is continuous with the execution step of the maximum end node, it can be determined that there is no target execution step.
[0129] If the execution steps of the starting node are not continuous with the execution steps of the maximum terminating node, the execution steps between the two execution steps can be used as the target execution steps, wherein the execution steps of the starting node, the target execution steps and the execution steps of the maximum terminating node are continuous execution steps, and the target execution steps include at least one.
[0130] For example, Figure 3 Taking the neural network model shown as an example, the starting node is node conv_1, and the maximum terminal node is node add_1, wherein the execution step of the starting node is execution step 1, the execution step of the maximum terminal node is execution step 3, and the target execution step is execution step 2.
[0131] For example, Figure 4 Taking the neural network model shown as an example, for the first skip connection, the starting node is node conv_4, and the maximum terminating node is node add_2, wherein the execution step of the starting node is execution step 1, and the execution step of the maximum terminating node is execution step 3, then the target execution step is execution step 2; for the second skip connection, the starting node is node conv_5, and the maximum terminating node is node add_3, wherein the execution step of the starting node is execution step 2, and the execution step of the maximum terminating node is execution step 5, then the target execution steps are execution steps 3 and execution step 4.
[0132] In the memory management method provided in the embodiment of the present application, node pairs can be screened from all node pairs of the target neural network model to obtain the maximum termination node corresponding to the same starting node; the target execution steps are determined according to the execution steps of the starting node and the execution steps of the maximum termination node. Among them, the target execution steps can be accurately determined through the execution steps of the starting node and the execution steps of the maximum termination node, and then the memory occupied by the target execution steps can be more accurately determined.
[0133] In one embodiment, the memory occupied by the target neural network model at each execution step is determined based on the memory occupied by the computing nodes running the target neural network model at each execution step, including: determining the memory occupied by the target neural network model at the target execution step based on the memory occupied by the initial computing node of the target execution step at runtime and the memory occupied by the starting node of the jump connection at runtime.
[0134] It should be noted that in the process of calculating the memory occupied by the target neural network model in the target execution step, the sum of the memory occupied by the initial computing node of the target execution step at runtime and the memory occupied by the starting node of the jump connection at runtime can be used as the memory occupied by the target execution step.
[0135] In the process of determining the memory occupied by the target execution step, the actual running computing node corresponding to each target execution step can be determined by expanding the neural network model.
[0136] It should be noted that for the step of skip connection, there are multiple execution branches, such as Figure 1 Taking the neural network model shown in the figure as an example, node 2 is the jump connection of the starting node. The calculation of nodes 3 and 4 is realized through the output of node 2 respectively. Since node 5 also needs the output of node 2 during the calculation process, the calculation of node 2 will not be completed during the calculation of nodes 3 and 4. Figure 1 The skip connection steps shown are expanded to obtain the following results:
[0137] Fig.10 This is a schematic diagram of the result of expanding the neural network model provided in the embodiment of the present application, please refer to Fig.10 , because node 2 is also running during the calculation of nodes 3 and 4, the actual running computing nodes in step 3 corresponding to nodes 3 and 4 include node 2 which keeps running. The expanded neural network model is as follows Fig.10 As shown, the computing nodes corresponding to executing step 3 include: node 2, node 3 and node 4.
[0138] That is to say, the starting node of the skip connection will continue to occupy memory until the skip connection ends. During the memory calculation process, the sum of the memory occupied by the initial calculation node of the target execution step at runtime and the memory occupied by the starting node of the skip connection at runtime can be used as the memory occupied by the target execution step.
[0139] In the memory management method provided in the embodiment of the present application, the memory occupied by the target neural network model in the target execution step can be determined according to the memory occupied by the initial computing node of the target execution step at run time and the memory occupied by the starting node of the jump connection at run time. For a target execution step with a jump connection, since the starting node of the jump connection is always in a running state during the jump connection process, the memory occupied by the initial computing node of the target execution step at run time and the memory occupied by the starting node of the jump connection at run time can be combined to more accurately determine the memory occupied by the target neural network model in the target execution step.
[0140] Another feasible implementation of the memory management method provided in the embodiment of the present application is explained below.
[0141] Fig.11 Another flowchart of the memory management method provided in the embodiment of the present application is shown in FIG. Fig.11 Before determining the memory occupied by the target neural network model according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, the method further includes:
[0142] S1110: Determine the memory occupied by each computing node according to the type of each computing node and the size of input and output data of each computing node.
[0143] It should be noted that before determining the memory occupied by the target neural network model, the memory occupied by each computing node during runtime can be determined.
[0144] The type of the computing node may refer to the type of function actually run by the computing node, for example, a computing node for convolution calculation, a computing node for accumulation calculation, etc.
[0145] The size of the input and output data of a computing node may be the data size of the input data and output data of the computing node. For example, for image data, the size of the data may be determined based on the number of pixels and resolution of the image; for audio data, the size of the audio data may be determined based on the duration of the audio data and the number of audio signals included in each frame; for other different types of data, the size may also be determined in a corresponding manner.
[0146] After the type of each computing node and the size of input and output data of each computing node are obtained, the occupied memory of each computing node can be determined according to the type of each computing node and the size of input and output data of each computing node.
[0147] For example, if a computing node is a computing node for accumulation calculation, the memory required in the accumulation process can be determined according to the type of computing node, and the memory occupied by the input and output data of the computing node can be determined according to the size of the data.
[0148] For example, assuming that the cumulative calculation requires 1 mb of computing space, and the size of the input and output data of each computing node requires 10 mb of computing space, it can be determined that the memory occupied by the computing node is the sum of the two, that is, 11 mb.
[0149] It should be noted that the above process of determining each computing node is only one feasible example, and other methods may be used in actual implementation, which is not specifically limited here.
[0150] In the memory management method provided in the embodiment of the present application, the occupied memory of each computing node can be determined according to the type of each computing node and the size of the input and output data of each computing node. Among them, by pre-calculating the occupied memory of each computing node, the memory occupied by the target neural network model in each step can be more quickly performed, thereby improving the efficiency of memory management.
[0151] Another feasible implementation of the memory management method provided in the embodiments of the present application is explained below.
[0152] Fig.12 This is another flowchart of the memory management method provided in the embodiment of the present application. Please refer to Fig.12 , after determining the memory occupied by the target neural network model at each execution step according to the memory occupied by the computing nodes running the target neural network model at each execution step, the method further includes:
[0153] S1210: Determine the steps to be optimized of the target neural network model according to the memory occupied by the target neural network model at each execution step.
[0154] Among them, the steps to be optimized are steps whose occupied memory is greater than or equal to a preset memory threshold.
[0155] It should be noted that after determining the memory occupied by each execution step, the steps to be optimized can be determined from these execution steps, wherein the steps to be optimized refer to the steps that currently occupy a large amount of memory and need to reduce the memory occupation through optimization.
[0156] In one embodiment, a preset memory threshold may be set, and steps that occupy memory less than the preset memory threshold are treated as steps that do not need to be optimized, and steps that occupy memory greater than or equal to the preset memory threshold are treated as steps to be optimized.
[0157] Optionally, after the steps to be optimized are obtained, the steps to be optimized may be fed back to the corresponding level of the electronic device, for example, the steps to be optimized may be fed back to a display interface to display the steps that can be optimized.
[0158] It should be noted that there may be an application in the electronic device for monitoring the real-time memory of the target neural network model. The application can display the memory occupied by the target neural network model in each execution step and can output the corresponding steps to be optimized, thereby outputting the corresponding steps to be optimized to the user of the application.
[0159] In the memory management method provided in the embodiment of the present application, the steps to be optimized of the target neural network model can be determined according to the memory occupied by the target neural network model in each execution step. By outputting the steps to be optimized, the user can be reminded to optimize the corresponding steps of the target neural network model.
[0160] The following is an explanation of one feasible implementation method for displaying the real-time memory usage of the target neural network model in the display interface of an electronic device.
[0161] Fig.13 Please refer to the calculation result diagram of the memory usage of each step provided in the embodiment of this application. Fig.13 , Fig.13 What is shown is a schematic diagram of the memory changes of the target neural network model within a preset time period, wherein the horizontal axis represents time, and each time interval may correspond to one of the execution steps of the target neural network model; the vertical axis represents the size of the occupied memory.
[0162] This schematic diagram can be used to determine the changes in the memory occupied by the target neural network model during the execution of each execution step.
[0163] It should be understood that, although the steps in the above-mentioned flowcharts are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0164] Based on the foregoing embodiments, an embodiment of the present application provides a memory management device, which includes the modules included and the units included in the modules, and can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0165] Fig.14 For a schematic diagram of the structure of the memory management device provided in the embodiment of the present application, please refer to Fig.14 In another aspect of the embodiment of the present application, a memory management device is provided, which is applied to an electronic device, wherein the electronic device runs a target neural network model, wherein the target neural network model includes a plurality of computing nodes, and the plurality of computing nodes run in a target execution order to implement the function of the target neural network model, wherein the device includes: a node determination module 1410, a memory calculation module 1420;
[0166] The node determination module 1410 is used to determine the initial computing node included in each execution step according to the target execution order;
[0167] The memory calculation module 1420 is used to determine the memory occupied by the target neural network model when a jump connection exists in the target neural network model, based on the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial calculation node included in each execution step at runtime, wherein the execution steps of the starting node and the ending node of the jump connection are discontinuous, and the calculation nodes of the target neural network model running in the target execution step include: the starting node of the jump connection and the initial calculation node of the target execution step; the target execution step is an execution step located between the starting execution step corresponding to the starting node and the ending execution step corresponding to the ending node.
[0168] In one embodiment, the memory computing module 1420 is specifically used to determine the computing node on which the target network model runs in executing the target execution step based on the starting node of the jump connection and the initial computing node included in the target execution step; and determine the memory occupied by the target neural network model in each execution step based on the memory occupied by the computing node on which the target neural network model runs in each execution step, each execution step including the target execution step.
[0169] In one embodiment, the memory computing module 1420 is also used to determine the next computing node corresponding to each computing node in the target neural network model in the target execution order; when the target computing node corresponds to at least two next computing nodes, and the execution steps corresponding to at least two next computing nodes are different, it is determined that there is a jump connection in the target neural network model, wherein the target computing node is the starting node of the jump connection.
[0170] In one embodiment, the memory computing module 1420 is specifically used to obtain all node pairs in the target neural network model, and the node pairs are used to represent any two computing nodes with a connection relationship in the target neural network model, and each node pair includes: a starting node and an ending node.
[0171] In one embodiment, the memory computing module 1420 is specifically used to determine whether a jump connection exists in the target neural network model when, in each node pair, the same starting node corresponds to at least two different terminating nodes and the execution steps of at least two different terminating nodes are different.
[0172] In one embodiment, the memory computing module 1420 is also used to screen node pairs from all node pairs of the target neural network model to obtain the maximum termination node corresponding to the same starting node, and the maximum termination node is the termination node with the latest execution steps in the target neural network model; according to the execution steps of the starting node and the execution steps of the maximum termination node, the target execution steps are determined, wherein the execution steps of the starting node, the target execution steps and the execution steps of the maximum termination node are continuous execution steps, and the target execution steps include at least one.
[0173] In one embodiment, the memory calculation module 1420 is specifically used to determine the memory occupied by the target neural network model at the target execution step based on the memory occupied by the initial computing node of the target execution step at runtime and the memory occupied by the starting node of the jump connection at runtime.
[0174] In one embodiment, the memory calculation module 1420 is further used to determine the occupied memory of each computing node according to the type of each computing node and the input and output of each computing node.
[0175] In one embodiment, the memory calculation module 1420 is also used to determine the steps to be optimized of the target neural network model based on the memory occupied by the target neural network model at each execution step, and the steps to be optimized are steps whose occupied memory is greater than or equal to a preset memory threshold.
[0176] In the memory management device provided in the embodiment of the present application, the initial computing node included in each execution step can be determined according to the target execution order; in the case where the target neural network model has a jump connection, the memory occupied by the target neural network model can be determined according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime. Among them, the target execution step can be an execution step located between the starting execution step corresponding to the starting node and the terminating execution step corresponding to the terminating node. In the process of calculating the occupied memory, the calculation can be performed according to the computing node running at the target execution step of the target neural network model, that is, the calculation of the occupied memory corresponding to the target execution step can be realized according to the starting node of the jump connection and the initial computing node of the target execution step, so that the memory occupied by the target neural network model in each execution step can be accurately determined.
[0177] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0178] It should be noted that in the embodiments of this application Fig.14 The division of modules in the memory management device shown is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.
[0179] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0180] Fig.15 For a schematic diagram of the structure of the computer device provided in the embodiment of the present application, please refer to Fig.15 , an embodiment of the present application provides a computer device, which may be the above-mentioned electronic device, and its internal structure diagram may be as shown in Fig.15 As shown. The computer device includes a processor 1520, a memory and a network interface 1540 connected via a system bus 1510. Among them, the processor 1520 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 1531 and an internal memory 1532. The non-volatile storage medium 1531 stores an operating system, a computer program and a database. The internal memory 1532 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium 1531. The database of the computer device is used to store data. The network interface 1540 of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor 1520, the above method is implemented.
[0181] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.
[0182] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0183] Those skilled in the art will understand that Fig.15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0184] In one embodiment, the memory management device provided by the present application can be implemented in the form of a computer program. The computer program can be Fig.15 The computer device shown in the figure can be run. The memory of the computer device can store various program modules constituting the above-mentioned device. The computer program composed of various program modules enables the processor to execute the steps of the method of each embodiment of the present application described in this specification.
[0185] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0186] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.
[0187] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0188] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0189] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0190] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0191] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0192] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0193] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0194] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0195] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0196] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0197] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A memory management method, characterized in that: Applied to an electronic device, the electronic device runs a target neural network model, the target neural network model includes a plurality of computing nodes, the plurality of computing nodes run in a target execution order to implement the function of the target neural network model, the method includes: Determining the initial computing nodes included in each execution step according to the target execution order; In the case where the target neural network model has a skip connection, the memory occupied by the target neural network model is determined according to the memory occupied by the starting node of the skip connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, wherein the execution steps of the starting node and the ending node of the skip connection are discontinuous, and the computing nodes of the target neural network model running in the target execution step include: the starting node of the skip connection and the initial computing node of the target execution step; and the target execution step is an execution step located between the starting execution step corresponding to the starting node and the ending execution step corresponding to the ending node.
2. The method according to claim 1, characterized in that Determining the memory occupied by the target neural network model according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime includes: Determine the computing node on which the target network model runs when executing the target execution step according to the starting node of the jump connection and the initial computing node included in the target execution step; The memory occupied by the target neural network model in each execution step is determined according to the memory occupied by the computing nodes running the target neural network model in each execution step, and each execution step includes the target execution step.
3. The method according to claim 1 or 2, characterized in that: Before determining the memory occupied by the target neural network model according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, the method further includes: Determining the next computational node corresponding to each computational node in the target neural network model in the target execution order; When a target computing node corresponds to at least two next computing nodes and the execution steps corresponding to the at least two next computing nodes are different, it is determined that a jump connection exists in the target neural network model, wherein the target computing node is the starting node of the jump connection.
4. The method according to claim 3, characterized in that The determining the next computing node corresponding to each computing node in the target neural network model in the target execution order includes: All node pairs in the target neural network model are obtained, where the node pairs are used to represent any two computing nodes having a connection relationship in the target neural network model, and each of the node pairs includes: a starting node and an ending node.
5. The method according to claim 4, characterized in that The step of determining that the target neural network model has a skip connection when the target computing node corresponds to at least two next computing nodes and the at least two next computing nodes correspond to different execution steps includes: In each of the node pairs, when the same starting node corresponds to at least two different terminating nodes and the execution steps of the at least two different terminating nodes are different, it is determined that the target neural network model has a skip connection.
6. The method according to claim 4, characterized in that Before determining the computing node on which the target network model is to be executed based on the starting node of the jump connection and the initial computing node included in the target execution step, the method further includes: Filter node pairs from all node pairs of the target neural network model to obtain the maximum termination node corresponding to the same starting node, wherein the maximum termination node is the termination node at the end of the execution step in the target neural network model; The target execution steps are determined according to the execution steps of the starting node and the execution steps of the maximum terminating node, wherein the execution steps of the starting node, the target execution steps and the execution steps of the maximum terminating node are continuous execution steps, and the target execution steps include at least one.
7. The method according to claim 2, characterized in that Determining the memory occupied by the target neural network model at each execution step according to the memory occupied by the computing nodes running the target neural network model at each execution step includes: The memory occupied by the target neural network model in the target execution step is determined according to the memory occupied by the initial computing node of the target execution step at runtime and the memory occupied by the starting node of the jump connection at runtime.
8. The method according to claim 1, characterized in that Before determining the memory occupied by the target neural network model according to the memory occupied by the starting node of the jump connection at runtime and the memory occupied by the initial computing node included in each execution step at runtime, the method further includes: The memory occupied by each computing node is determined according to the type of each computing node and the input and output of each computing node.
9. The method according to claim 1, characterized in that: After determining the memory occupied by the target neural network model at each execution step according to the memory occupied by the computing nodes running the target neural network model at each execution step, the method further comprises: The steps to be optimized of the target neural network model are determined according to the memory occupied by the target neural network model in each execution step, and the steps to be optimized are steps whose occupied memory is greater than or equal to a preset memory threshold.
10. A memory management device, characterized in that: Applied to an electronic device, the electronic device runs a target neural network model, the target neural network model includes a plurality of computing nodes, the plurality of computing nodes run in a target execution order to implement the function of the target neural network model, the device includes: a node determination module, a memory calculation module; The node determination module is used to determine the initial computing node included in each execution step according to the target execution order; The memory calculation module is used to determine the memory occupied by the target neural network model when there is a jump connection in the target neural network model, based on the memory occupied by the starting node of the jump connection at run time and the memory occupied by the initial calculation node included in each execution step at run time, wherein the execution steps of the starting node and the ending node of the jump connection are discontinuous, and the calculation nodes of the target neural network model running in the target execution step include: the starting node of the jump connection and the initial calculation node of the target execution step; the target execution step is an execution step located between the starting execution step corresponding to the starting node and the ending execution step corresponding to the ending node.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method and device for obtaining multi-hop neighbor node
CN106712995A
Node storage method and system based on neural network, server and storage medium
CN111309265A
Neural network model calculation method, data processing method, electronic equipment and medium
CN113657584A
Neural network memory optimization method and device based on hardware accelerator
CN114298294A
Memory information determination method and device and electronic equipment
CN116932218A