Operation method of neural network model, chip, electronic equipment and storage medium

By determining and optimizing the storage space order of running branches in the neural network model, the storage space occupation problem caused by random execution order is solved, and the operation efficiency and storage space reuse is improved.

CN120144258APending Publication Date: 2025-06-13ARM TECH CHINA CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510308664.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When running a neural network model, random selection of execution order may cause some operators to actually occupy a large amount of storage space during operation, affecting efficiency.

Method used

By determining the storage space required for each running branch of the first operator to run from N second intermediate operators, and running each running branch in the order of the storage space from small to large, to obtain N running results.

Benefits of technology

By optimizing the execution sequence, the speed of determining the operation result of the second target operator is reduced and the multiplexing efficiency of the storage space is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144258A_ABST
    Figure CN120144258A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses an operation method of a neural network model, a chip, electronic equipment and a storage medium. In the method, the electronic equipment firstly obtains a neural network model to be operated; the model comprises a first operator, a second operator and K intermediate operators. The first operator directly depends on M first intermediate operators in the K intermediate operators, and the second operator directly depends on N second intermediate operators in the K intermediate operators. A first operation branch which operates from the first operator to a first target operator in the N second intermediate operators and a second operation branch which operates from the first operator to the second target operator comprise at least one same intermediate operator. And determining a storage space required when each operation branch is operated in N operation branches which are respectively operated from the first operator to the N second intermediate operators, and operating the N operation branches based on the sequence of the storage spaces from small to large. According to the method, the storage space actually occupied by the second operation branch during operation can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method for running a neural network model, a chip, an electronic device, and a storage medium. Background Art

[0002] With the continuous development of computer technologies, neural network models are increasingly widely used in performing various tasks. For example, neural network models can be used to perform image processing tasks, natural language processing tasks, autonomous driving tasks, and so on.

[0003] The model structure of a neural network model usually includes multiple operators, and there is a certain dependency relationship among the multiple operators. For example, the operation of one operator depends on the operation result of another operator, that is, the operation result of the other operator can be used as at least part of the input data of the one operator.

[0004] It can be understood that there are multiple possible execution orders for multiple operators without a dependency relationship. Taking multiple operators including operator a and operator b as an example, the electronic device can first run operator a or first run operator b. However, if a certain execution order is randomly selected to run the multiple operators, it may cause a relatively large storage space actually occupied by some operators during the operation process. Summary of the Invention

[0005] Embodiments of this application provide a method for running a neural network model, a chip, an electronic device, and a storage medium.

[0006] In a first aspect, this application provides a method for running a neural network model. The method is applied to an electronic device and includes: obtaining a neural network model to be run, where the neural network model includes multiple operators, the multiple operators include a first operator, a second operator, and K intermediate operators between the first operator and the second operator, M first intermediate operators among the K intermediate operators directly depend on the first operator, the second operator directly depends on N second intermediate operators among the K intermediate operators, there is no dependency relationship between any two of the N second intermediate operators, and the M first intermediate operators and the N second intermediate operators are different. The first operation branch from the first operator to the first target operator among the N second intermediate operators and the second operation branch from the first operator to the second target operator among the N second intermediate operators include at least one same intermediate operator, K is a positive integer greater than 3, M and N are positive integers greater than 1; determining, for each of the N operation branches from the first operator to the N second intermediate operators respectively, the storage space required when each operation branch runs; and running the N operation branches in the order of the storage spaces corresponding to the N operation branches from small to large to obtain N operation results corresponding to the N operation branches one by one.

[0007] It can be understood that since the first execution branch and the second execution branch include at least one identical intermediate operator, if the storage space required by the second execution branch is greater than that required by the first execution branch, it indicates that the intermediate operator of the second execution branch is more likely to depend on the intermediate operator of the first execution branch. If the first execution branch is executed first and then the second execution branch is executed, since the identical intermediate operator has been executed when the first execution branch is executed, the actual storage space occupied when the second execution branch is executed will be less, thereby improving the determination speed of the operation result of the second target operator.

[0008] In a possible implementation of the first aspect, the first execution branch includes: among the operators that must be executed during the process from the first operator to the first target operator, the operators that have a direct or indirect dependency relationship with the first operator; the second execution branch includes: among the operators that must be executed during the process from the first operator to the second target operator, the operators that have a direct or indirect dependency relationship with the first operator.

[0009] It can be understood that the operators that must be executed during the process from the first operator to the first target operator include the first operator and the first target operator. Taking the operator A in Figure 1 as the operator A and the first target operator as the operator a4 as an example, the operators that must be executed during the process from the first operator to the first target operator include: operator A, operator a1, operator a2, operator a3, operator a4. Then, among the operators that must be executed during the process from the first operator to the first target operator, the operators that have a direct or indirect dependency relationship with the first operator include: operator a1, operator a2, operator a3, operator a4. Among them, operator a1 has a direct dependency relationship with operator A, and operator a2, operator a3, and operator a4 have an indirect dependency relationship with operator A.

[0010] In a possible implementation of the first aspect, the storage space required when the third execution branch among the N execution branches is executed includes: the storage space occupied in the memory and / or cache of the electronic device when the third execution branch is executed.

[0011] It can be understood that when the third execution branch among the N execution branches is executed, it may need to occupy the storage space in the memory of the electronic device or may occupy the storage space in the cache. Therefore, the electronic device can determine the storage space occupied in the memory and / or (high-speed) cache of the electronic device when the third execution branch is executed before the third execution branch is executed.

[0012] Among them, the third running branch can be any one of the N running branches. That is to say, the storage space required for each of the N running branches during operation includes: the storage space occupied by each running branch in the memory and / or cache of the electronic device during operation.

[0013] In a possible implementation of the first aspect, the storage space required for the third running branch during operation includes: the sum of the data volumes of the input data and / or output data of each operator in the third running branch.

[0014] It can be understood that the size relationship between the data volumes of the input data and output data of each operator is determined based on the structure of the corresponding operator. Therefore, when the structure of the operator is determined, the size relationship between the input data and output data of the operator can be determined. At this time, if the electronic device can determine the data volume of the input data of the operator (for example, taking the output data of the operator directly depended on by a certain operator as the input data of the certain operator), the electronic device can then determine the data volume of the output data of the operator based on the data volume of the input data and the size relationship between the input data and output data, and further can use the data volume of the input data and / or output data as the storage space required for each operator during operation.

[0015] This application does not limit the type of the data volume of the input data and / or output data. For example, the data volume of the input data and / or output data can be reflected in the form of the number of bits, or in the form of the number of bytes, or in any form that can characterize the data volume of the input data and / or output data.

[0016] It can be understood that the electronic device needs to obtain the input data during the operation of each operator, and also needs to store the obtained output data after the operation is completed. Therefore, the method of using the data volume of the input data and / or output data of each operator as the storage space required for the operator during operation can make the determination of the storage space required for the operator during operation more accurate, and further make the determination of the storage space required for each of the N running branches during operation more accurate.

[0017] In a possible implementation of the first aspect, the method further includes: after running the N running branches, running a second operator based on the N running results to obtain the running result of the second operator.

[0018] For example, the electronic device can be at Figure 1Based on the shown model structure, based on the operation results of the operation branches corresponding to operator a4, operator b4, operator c4, and operator d4, operator B (as an example of the second operator) is run to obtain the operation result of operator B. Thus far, Figure 1 All the operators in the shown model structure have been executed.

[0019] In a possible implementation of the first aspect, the multiple operators include one or more of the following operators: arithmetic operator, convolution operator, pooling operator, normalization operator, fully connected operator, activation operator, recurrent operator, attention operator, tensor transformation operator, tensor splicing and splitting operator, loss function operator, upsampling and interpolation operator.

[0020] In a possible implementation of the first aspect, the neural network model includes one or more of a deep learning model, a generative model, a machine learning model, and a reinforcement learning model; and the neural network model is used to process one or more of image data, text data, audio data, vehicle state data, and time series data.

[0021] In a second aspect, the present application provides a chip, which is used to implement the operation method of the neural network model in the first aspect and any possible implementation of the first aspect.

[0022] In a third aspect, the present application provides an electronic device, which includes the chip related to the second aspect, and one or more memories; one or more programs are stored in the one or more memories, and when the one or more programs are executed by the chip, the electronic device executes the operation method of the neural network model in the first aspect and any possible implementation of the first aspect.

[0023] In a fourth aspect, the present application provides a readable storage medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device executes the operation method of the neural network model in the first aspect and any possible implementation of the first aspect.

[0024] Among them, the beneficial effects of the second aspect to the fourth aspect can refer to the beneficial effects of the first aspect and any possible implementation of the first aspect, which will not be elaborated here. Description of the Drawings

[0025] Figure 1 According to some embodiments of the present application, a partial structural schematic diagram of a neural network model is shown;

[0026] Figure 2 According to some embodiments of the present application, a flowchart of an operation method of a neural network model is shown;

[0027] Figure 3 According to some embodiments of the present application, a schematic flow chart of a method for determining the storage space required during the operation of N running branches is shown;

[0028] Figure 4 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown. Specific embodiments

[0029] The illustrative embodiments of the present application include, but are not limited to, methods for running a neural network model, chips, electronic devices, and storage media.

[0030] To facilitate understanding of the method for running a neural network model provided by the present application, the following first explains the proprietary terms involved in the embodiments of the present application.

[0031] Neural network model: A computational model that simulates the human nervous system obtained through computer algorithms and data training. Through the neural network model, an electronic device can complete a series of tasks including image recognition, speech recognition, natural language processing, machine translation, etc. The model structure of the neural network model is usually determined before the training process so that the neural network model can perform specific tasks.

[0032] Operator: The basic building block of a neural network model, which can serve as a computational node in the neural network model, used to perform basic mathematical operations or data transformations, and to achieve data transmission. In the present application, an operator can also be referred to as a node, layer, etc.

[0033] There is a certain dependency relationship among multiple operators of the neural network model, such as including direct dependency relationships and indirect dependency relationships. Multiple operators with a certain dependency relationship can constitute the model structure of the neural network model. Among them, a partial model structure diagram of the neural network model can be seen in Figure 1 . As Figure 1 shown, the neural network model includes, but is not limited to, eighteen operators, namely operator A, operators a1 - a4, operators b1 - b4, operators c1 - c4, operators d1 - d4, and operator B.

[0034] Among them, Figure 1 the direction indicated by the arrows between the operators in Figure 1 represents the dependency relationship between the operators or the data flow of the operation results of the operators, and does not represent that there are

[0035] the arrows shown in Figure 1Operator A in it, and the other operator is Figure 1 Take operator a1 in it as an example, such as Figure 1 As shown, operator a1 directly depends on operator A, and the output data of operator A is used as the input data of operator a1.

[0036] It can be understood that if there is an indirect dependence relationship between one operator and another operator, and the other operator indirectly depends on the one operator, then the input data of the other operator includes the output data obtained by processing the output data of the one operator through one or more other operators. Among them, the input data of at least one of the one or more other operators includes the output data of the one operator.

[0037] Take the one operator as Figure 1 Operator a1 in it, and the other operator is Figure 1 Take operator b3 in it as an example, such as Figure 1 As shown, the input data of operator b3 includes the output data obtained by processing the output data of operator a1 through operator a2 and operator b2 (operator a2 and operator b2 are taken as examples of one or more other operators). Among them, the input data of operator a2 and operator b2 (operator a2 and operator b2 are taken as examples of at least one other operator) includes the output data of operator a1.

[0038] Take the one operator as Figure 1 Operator a1 in it, and the other operator is Figure 1 Take operator b4 in it as an example, such as Figure 1 As shown, the input data of operator b4 includes the output data obtained by processing the output data of operator a1 through operator a2 - a3 and operator b2 - b3 (operator a2 - a3 and operator b2 - b3 are taken as examples of one or more other operators). Among them, the input data of operator a2 and operator b2 (operator a2 and operator b2 are taken as examples of at least one other operator) in operator a2 - a3 and operator b2 - b3 includes the output data of operator a1.

[0039] Common operator types include: arithmetic operator, convolution operator, pooling operator, normalization operator, fully - connected operator, activation operator, recurrent operator, attention operator, tensor deformation operator, tensor splicing and splitting operator, loss function operator, up - sampling and interpolation operator, etc.

[0040] Different types of operators can be used to perform different computational tasks. For example, arithmetic operators can be used to perform basic mathematical operations such as addition, subtraction, multiplication, and division. Convolution operators can be used to perform convolution operations based on convolution kernels of different sizes and shapes to extract features of data in different dimensions. Pooling operators can be used to reduce the dimension of data, retain important features, and improve computational efficiency. Normalization operators can be used to convert data with different dimensions or ranges into a unified dimensionless form, so that the data has the same dimension and range, and improve the standardization degree of the data. Fully connected operators can be used to fuse the features of the previous layer and improve the data understanding ability of the neural network model.

[0041] Activation operators are used to introduce non-linear transformations so that the neural network model can learn complex patterns. Recurrent operators are used to process sequential data and capture dependencies between time steps. Attention operators are used to dynamically assign weights to each operator. Tensor reshaping operators are used to adjust the data dimension to fit the network structure. Tensor concatenation and splitting operators are used to merge or split data streams to achieve information fusion or parallel processing. Loss function operators are used to quantify the difference between the model prediction value and the true value and perform parameter updates. Upsampling and interpolation operators are used to enlarge the size of the feature map and restore spatial details.

[0042] Next, continue to combine Figure 1 , to introduce the technical solution of this application.

[0043] It should be noted that the neural network model in this application can be any artificial intelligence model, including but not limited to deep learning models (such as convolutional neural network models, recurrent neural network models), generative models (such as diffusion models), machine learning models (such as logistic regression models, decision tree models), reinforcement learning models, transfer learning models, and so on.

[0044] The neural network model in this application has multiple different task processing functions and can process multiple different types of data. For example, the neural network model can process image data to perform tasks such as image recognition and image generation; it can process text data to perform natural language processing tasks (such as machine translation, dialogue generation, poetry creation); it can process audio data to perform tasks such as speech recognition, speech synthesis, music classification, and music synthesis; it can process the state data of vehicles to perform autonomous driving tasks; it can also process time series data to perform prediction and classification tasks.

[0045] It can be understood that an electronic device usually runs a neural network model based on a neural processing unit (NPU). For the case where an electronic device runs a neural network model based on only one neural network processor, the neural network processor can only run one operator at a time. However, there are multiple possible execution orders for operators that have no dependencies. If a random execution order is selected from the multiple possible execution orders to run multiple operators, it may cause some operators to actually occupy a relatively large storage space during operation.

[0046] Continuing with Figure 1 the model structure shown as an example, operator a1, operator b1, operator c1, and operator d1 directly depend on operator A respectively, operator B directly depends on operator a4, operator b4, operator c4, and operator d4 respectively, and there are no dependencies among operator a4, operator b4, operator c4, and operator d4. Among them, Figure 1 the determination methods of the direct dependencies and indirect dependencies among the operators shown have been described above, and will not be elaborated one by one.

[0047] In Figure 1 the model structure shown, there are multiple possible execution orders for operator a4, operator b4, operator c4, and operator d4. For example, the execution order is: operator a4 - operator b4 - operator c4 - operator d4 (hereinafter referred to as execution order_1). For another example, the execution order is: operator b4 - operator a4 - operator c4 - operator d4 (hereinafter referred to as execution order_2).

[0048] In the case of running each operator based on execution order_1, since operator a4 directly depends on operator a3, operator a3 directly depends on operator a2, and operator a2 directly depends on operator a1, when running operator a4, it is necessary to run operator a1, operator a2, and operator a3 all. Taking Figure 1 the case where the storage space occupied by each operator during its own operation is the same and is one unit of storage space as an example, operator a4 actually occupies a total of 4 units of storage space when running.

[0049] In addition, since operator b4 directly depends on operator a3 and operator b3, operator b3 directly depends on operator a2 and operator b2, operator b2 directly depends on operator a1 and operator b1, and operator a1, operator a2, and operator a3 have already been run when operator a4 is running, when running operator b4, there is no need for additional storage space to run operator a1, operator a2, and operator a3. That is to say, operator b4 actually occupies a total of 4 units of storage space when running.

[0050] Similarly, operator c4 and operator d4 also actually occupy a total of 4 units of storage space when running respectively.

[0051] It can be understood that since operator A needs to run when any operator runs, the storage space of operator A is not included in the above calculation of the storage space.

[0052] In the case of running each operator based on execution order _2, based on the calculation method of the storage space described above, when operator b4 runs, it actually occupies a total of 7 units of storage space; when operator a4 runs, it actually occupies a total of 1 unit of storage space; when operator c4 runs, it actually occupies a total of 4 units of storage space; when operator d4 runs, it actually occupies a total of 4 units of storage space.

[0053] It can be seen that compared with execution order _1, when running each operator based on execution order _2, the actual storage space occupied when operator b4 runs is larger, and since the number of operators to be executed is larger, the determination speed of the running result of operator b4 is also slower, reducing the running efficiency of operator b4.

[0054] In addition, in the case of running each operator based on execution order _1, since operator a4 runs first, after operator a4 finishes running, part of the storage space of operator a1, operator a2, operator a3, and operator a4 can be released. Taking the running time of each operator as the same and all being one unit of time as an example, after 4 units of time, the electronic device can release part of the storage space for the first time.

[0055] In the case of running each operator based on execution order _2, since operator b4 runs first, after operator b4 finishes running, part of the storage space of operator a1 - a3 and operator b1 - b4 can be released. Taking the time required for each operator to run as one unit of time as an example, after 7 units of time, the electronic device can release part of the storage space for the first time.

[0056] Compared with execution order _1, when running each operator based on execution order _2, the first release time of the storage space is later, which is not conducive to the reuse of the storage space.

[0057] For the sake of convenience of description, hereinafter, among the operators that must be executed in the process of running from one operator to another operator, the operators that have a direct or indirect dependency relationship with this one operator are called the running branches from this one operator to that another operator. For example, for Figure 1 the model structure shown, operator A can be this one operator, and operator a4 can be that another operator. In this case, the running branches from operator A to operator a4 include: operator a1 - a4.

[0058] For another example, for Figure 1For the model structure shown, operator A can be one operator, and operator b4 can be the other operator. In this case, the running branch from operator A to operator b4 includes: operators a1 - a3 and operators b1 - b4.

[0059] Similarly, the running branch from operator A to operator c4 includes: operators a1 - a2, operators b1 - b3, and operators c1 - c4. The running branch from operator A to operator d4 includes: operator a1, operators b1 - b2, operators c1 - c3, and operators d1 - d4.

[0060] Therefore, in order to reduce the actual storage space occupied during the running of some running branches, the present application provides a method for running a neural network model, which is used to run a neural network model including at least a first operator, a second operator, and K intermediate operators between the first operator and the second operator. Among them, M first intermediate operators among the K intermediate operators directly depend on the first operator (that is, there is a direct dependence relationship between the first operator and the M first intermediate operators), and the second operator directly depends on N second intermediate operators among the K intermediate operators (that is, there is a direct dependence relationship between the second operator and the N second intermediate operators, and there is no dependence relationship between any two of the N second intermediate operators). The N second intermediate operators include a first target operator and a second target operator. The first running branch from the first operator to the first target operator and the second running branch from the first operator to the second target operator include at least one same intermediate operator. Among them, K is a positive integer greater than 3, and both M and N are positive integers greater than 1.

[0061] During the process of running the neural network model, the electronic device can first determine N running branches from the first operator to the N second intermediate operators respectively, and determine the storage space required during the running of each running branch. Then, the electronic device can run each running branch in the order of the storage space of the N running branches from small to large, and obtain N running results corresponding to the N running branches one by one.

[0062] Taking the case where the storage space required for the second running branch is greater than that required for the first running branch as an example, the electronic device can first run the first running branch and then run the second running branch.

[0063] It can be understood that since the first running branch and the second running branch include at least one same intermediate operator, if the storage space required for the second running branch is greater than that required for the first running branch, it means that the intermediate operators of the second running branch are more likely to depend on the intermediate operators of the first running branch. If the first running branch is run first and then the second running branch, since the same intermediate operators have been run during the running of the first running branch, the actual storage space occupied during the running of the second running branch will be less.

[0064] Continuing with Figure 1 the model structure shown as an example, if the storage space required for each operator during operation is the same and is one unit of storage space, taking the first target operator as operator a4 and the second target operator as operator b4 as an example, the storage space required for the first running branch from operator A to operator a4 during operation is 4 units of storage space, and the storage space required for the second running branch from operator A to operator b4 during operation is 7 units of storage space.

[0065] If the method provided in this application is adopted, the electronic device can first run the first running branch from operator A to operator a4, and then run the second running branch from operator A to operator b4.

[0066] It can be understood that since there are the same intermediate operators in the first running branch and the second running branch, therefore, if the first running branch is run first, then when the second running branch is run, the same intermediate operators (for example, operators a1 - a3) no longer require additional storage space for operation. That is to say, if the method provided in this application is adopted, the second running branch actually only occupies 4 units of storage space during operation (that is, the storage space corresponding to operators b1 - b4).

[0067] Therefore, this method can reduce the actual storage space occupied during the operation of the second running branch and improve the determination speed of the operation result of the second target operator.

[0068] In addition, in the method provided in this application, after the first running branch finishes running, a part of the storage space corresponding to the first running branch (such as a part of the storage space of operators a1 - a4) can be released. For example, the electronic device can first release a part of the storage space after 4 unit times. For other execution orders, such as the execution order of first executing the second running branch, the electronic device can first release a part of the storage space only after operator b4 finishes running (such as 7 time units), and the first release time of the storage space is relatively late. Therefore, the method provided in this application can shorten the time for the electronic device to first release the storage space, which is beneficial to the reuse of the storage space.

[0069] Before introducing in detail the operation method of the neural network model involved in the embodiments of this application, first, the electronic device applicable to this method is described.

[0070] It can be understood that the method for running the neural network model provided by the embodiments of the present application is applicable to any electronic device capable of running a neural network model. The electronic device may include, but is not limited to, a mobile phone, a smart TV, a wearable device, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, a wireless device in industrial control, a wireless device in self-driving, a wireless device in remote medical surgery, a wireless device in smart grid, a wireless device in transportation safety, a wireless device in smart city, a wireless device in smart home, etc. The embodiments of the present application do not limit the type of the electronic device.

[0071] The following describes in detail the method for running the neural network model provided by the present application with reference to the accompanying drawings. Figure 2 A schematic flowchart of a method for running the neural network model provided by the present application is shown. Figure 2 The execution subject of each step of the shown process is an electronic device. For the sake of convenience of description, the execution subject of each step will not be repeatedly described hereinafter when introducing Figure 2 each step of the shown process. As Figure 2 shown, the method includes, but is not limited to, the following solutions:

[0072] S201: Obtain a neural network model to be run.

[0073] The embodiments of the present application do not limit the type of the neural network model. The neural network model includes, but is not limited to, a deep learning model (for example, a convolutional neural network model, a recurrent neural network model), a generative model (for example, a diffusion model), a machine learning model (for example, a logistic regression model, a decision tree model), a reinforcement learning model, a transfer learning model, etc.

[0074] The neural network model in the present application can execute multiple tasks and process multiple data types. For example, the neural network model can process one or more of image data, text data, audio data, vehicle status data, and time series data. Among them, the data types processed by the neural network model and the types of tasks executed can be specifically referred to the relevant descriptions above, and will not be elaborated here.

[0075] In the embodiments of the present application, the neural network model includes multiple operators with certain dependencies. The multiple operators include a first operator, a second operator, and K intermediate operators between the first operator and the second operator. In some embodiments, the first operator may also be referred to as the starting operator, and the second operator may also be referred to as the ending operator.

[0076] It can be understood that the above-mentioned first operator, second operator, and K intermediate operators between the first operator and the second operator may be all the operators included in the neural network model of the present application, or may be part of the operators included in the neural network model. That is to say, the electronic device can run all the operators included in the neural network model based on the method provided in the embodiments of the present application; or, the electronic device can also run part of the operators included in the neural network model based on the method provided in the embodiments of the present application, and the operation of other operators except the above-mentioned part of the operators does not adopt the method provided in the embodiments of the present application.

[0077] In some embodiments, the first operator and the second operator may be virtual operators, that is, the first operator and the second operator do not perform specific operation tasks. In this case, the data volume (such as the number of bytes) of the input data (or called input tensor) and output data (or called output tensor) of the first operator is 0, and the data volume of the input data and output data of the second operator is also 0.

[0078] The embodiments of the present application do not limit the types of the multiple operators included in the neural network model, and the type of each operator can be adjusted according to the actual application scenario or actual task requirements. For example, the operators in the present application (such as the first operator, the second operator, the intermediate operator, etc.) may include arithmetic operators, convolution operators, pooling operators, normalization operators, fully connected operators, activation operators, cyclic operators, attention operators, tensor deformation operators, tensor splicing and splitting operators, loss function operators, upsampling and interpolation operators, and so on. The functions of different types of operators can be referred to the relevant descriptions above, and will not be elaborated here one by one.

[0079] In the neural network model, M first intermediate operators among the K intermediate operators directly depend on the first operator, and the second operator directly depends on N second intermediate operators among the K intermediate operators. There is no dependency relationship between any two of the N second intermediate operators, and the M first intermediate operators and the N second intermediate operators are different. Wherein, K is a positive integer greater than 3, M and N are positive integers greater than 1, and M and N may be the same or different, and the present application does not limit this.

[0080] It can be understood that the M first intermediate operators directly depend on the first operator, indicating that the operation result of the first operator can be used as at least part of the input data of the M first intermediate operators. Similarly, the second operator directly depends on the N second intermediate operators, indicating that the operation results of the N second intermediate operators can be used as at least part of the input data of the second operator.

[0081] In addition, the first operation branch from the first operator to the first target operator among the N second intermediate operators and the second operation branch from the first operator to the second target operator among the N second intermediate operators include at least one same intermediate operator.

[0082] Among them, the first operation branch includes: among the operators that must be executed in the process from the first operator to the first target operator, the operators that have a direct or indirect dependence relationship with the first operator.

[0083] Similarly, the second operation branch includes: among the operators that must be executed in the process from the first operator to the second target operator, the operators that have a direct or indirect dependence relationship with the first operator.

[0084] Taking at least part of the structure of the neural network model as Figure 1 the model structure shown as an example, in some embodiments, operator A can be used as the first operator, operator B can be used as the second operator, and the operators a1-a4, b1-b4, c1-c4, d1-d4 between operator A and operator B can be used as the N intermediate operators. At this time, N is equal to 16.

[0085] Operators a1, b1, c1, and d1 directly depend on operator A. Therefore, operators a1, a2, a3, and a4 can be used as the M first intermediate operators. At this time, M is equal to 4.

[0086] Operator B directly depends on operators a4, b4, c4, and d4. Therefore, operators a4, b4, c4, and d4 can be used as the N second intermediate operators. At this time, N is equal to 4.

[0087] Taking the first target operator as operator a4 as an example, the operation branch from operator A to operator a4 (i.e., the first operation branch) includes: operators a1-a4. Taking the second target operator as operator b4 as an example, the operation branch from operator A to operator b4 (i.e., the second operation branch) includes: operators a1-a3, b1-b4.

[0088] For the convenience of description, hereinafter, the operation branch from operator A to operator a4 will be expressed as the operation branch corresponding to operator a4, and the operation branch from operator A to operator b4 will be expressed as the operation branch corresponding to operator b4, and so on.

[0089] It can be understood that the embodiments of the present application do not limit the selection of the first operator and the second operator, as long as the selected first operator, second operator, and intermediate operator can meet the conditions described above.

[0090] Continuing with at least part of the structure of the neural network model as Figure 1 the model structure shown, in some other embodiments, operator b2 can also be used as the first operator, and operator B is still used as the second operator. In this case, operators b3 - b4, c3 - c4, and d4 between operator b2 and operator B can be used as N intermediate operators, where N is equal to 5 at this time.

[0091] Operators b3 and c3 directly depend on operator b2. Therefore, operators b3 and c3 can be used as M first intermediate operators, where M is equal to 2 at this time.

[0092] Operator B directly depends on operators b4, c4, and d4. Therefore, operators b4, c4, and d4 can be used as N second intermediate operators, where N is equal to 3 at this time.

[0093] Taking the first target operator as operator b4 as an example, the first running branch from operator b2 to operator b4 includes: operators b3 - b4. Taking the second target operator as operator c4 as an example, the second running branch from operator b2 to operator c4 includes: operators b3 - b4, c3 - c4.

[0094] In the case where the first operator and the second operator are other operators, the determination methods of the first intermediate operator, the second intermediate operator, the first target operator, the second target operator, the first running branch, and the second running branch are similar to the methods described above, and will not be elaborated one by one.

[0095] It can be understood that after the electronic device obtains the neural network model to be run, it can determine the model structure of the neural network model described above, such as the dependency relationship between multiple operators included in the neural network model, etc., to execute subsequent S202 and S203.

[0096] The embodiments of the present application do not limit the way for the electronic device to determine the model structure of the neural network model. For example, the electronic device can determine the dependency (or connection) relationship and other information between operators by parsing the serialized file of the neural network model (for example, the serialized file stores the dependency relationship between each operator included in the neural network model and other operators), and then determine the model structure of the neural network model.

[0097] For another example, an electronic device can parse each node in the static computational graph of a neural network model (the nodes in the static computational graph correspond one-to-one with the operators of the neural network model) and the edges (if there is an edge between any two nodes in the static computational graph, it indicates that there is a dependency relationship between the two nodes, and the direction of the edge indicates the data flow direction between the two nodes), determine information such as the dependency relationship between the operators, and further determine the model structure of the neural network model.

[0098] It can be understood that Figure 1 The model structure shown can be at least part of the structure of the neural network model. That is to say, in addition to including Figure 1 the model structure shown, the neural network model can also include other mutually dependent operators.

[0099] S202: Determine the storage space required for each of the N running branches when the first operator runs to the N second intermediate operators respectively.

[0100] It can be understood that since one second intermediate operator corresponds to one running branch, and one running branch corresponds to the storage space of one value, therefore, N second intermediate operators correspond to N running branches, and N running branches correspond to N storage spaces.

[0101] In the embodiments of the present application, the storage space required for each running branch when running can include: the space occupied by each running branch in the memory and / or cache of the electronic device when running.

[0102] In some embodiments, as Figure 3 shown, determining the storage space required for each of the N running branches when the first operator runs to the N second intermediate operators respectively includes the following S301 - S302.

[0103] S301: Determine the data volume of the input data and / or output data of each operator in each of the N running branches.

[0104] It can be understood that the size relationship between the input data and the output data of each operator of the neural network model is determined based on the structure of the corresponding operator. Therefore, when the structure of the operator is determined, the size relationship between the input data and the output data of the operator can be determined. At this time, if the electronic device can determine the data volume of the input data of the operator (for example, taking the output data of the operator directly depended on by a certain operator as the input data of the certain operator), then the electronic device can determine the data volume of the output data of the operator based on the data volume of the input data and the size relationship between the input data and the output data.

[0105] Taking the third execution branch among the N execution branches as an example, the electronic device may determine multiple operators included in the third execution branch, and then determine the data volume of the output data of each operator included in the third execution branch based on the data volume of the input data of each operator included in the third execution branch and the magnitude relationship between the data volumes of the input data and the output data.

[0106] This application does not limit the type of the data volume of the input data and / or the output data of each operator included in the third execution branch. For example, the data volume of the input data and / or the output data may be represented in the form of the number of bits, or in the form of the number of bytes, or in any form that can characterize the magnitude of the data volume of the input data and / or the output data.

[0107] This embodiment of the application also does not limit the manner of determining the magnitude of the data volume of the input data and / or the output data. Taking the data volume represented in the form of the number of bytes as an example, the electronic device may determine the total number of elements of the input data and / or the output data based on the dimensional distribution of the input data and / or the output data of a certain operator (for example, taking the product of the elements of each dimension as the total number of elements), and then take the product of the total number of elements and the number of bytes of each element as the data volume of the input data and / or the output data of the certain operator.

[0108] The above content only takes the third execution branch among the N execution branches as an example to describe the manner of determining the data volume of the input data and / or the output data of each operator in the third execution branch. However, it should be understood that among the N execution branches, the principle of determining the data volume of the input data and / or the output data of each operator in each execution branch is the same, and will not be elaborated one by one.

[0109] S302: Take the sum of the data volumes of the input data and / or the output data corresponding to each operator in each execution branch as the storage space required when the corresponding execution branch runs.

[0110] Continuing to take the third execution branch among the N execution branches as an example, it can be understood that the electronic device may take the sum of the data volumes of the input data and / or the output data corresponding to each operator in the third execution branch as the storage space required when the third execution branch runs.

[0111] Taking the third execution branch as the same as the first execution branch involved above as an example, the storage space required when the third execution branch runs is described below.

[0112] Taking the first operator and the second operator as Figure 1Taking the operator A and operator B in [[]] and the first target operator being operator a4 as an example, based on the foregoing, the first operation branch includes: operators a1 - a4. The storage space required when the first operation branch runs is equal to the sum of the storage spaces required when operators a1 - a4 run. If the storage space required for each operator to run is the same and is one unit of storage space, the first operation branch requires 4 units of storage space to run in this case.

[0113] In some other embodiments, taking the third operation branch being the same as the second operation branch involved above as an example, the storage space required when the third operation branch runs is described.

[0114] Taking the first operator and the second operator as Figure 1 the operator A and operator B in [[]] and the second target operator being operator b4 as an example, based on the foregoing, the second operation branch includes: operators a1 - a3, operators b1 - b4. The storage space required when the second operation branch runs is equal to the sum of the storage spaces required when operators a1 - a3 run and the storage spaces required when operators b1 - b4 run. If the storage space required for each operator to run is the same and is one unit of storage space, the storage space required for the second operation branch to run in this case is 7 units of storage space.

[0115] The above content only takes the third operation branch as the first operation branch and the third operation branch as the second operation branch as examples to describe the calculation method of the storage space required when the third operation branch runs. However, it should be understood that the determination principle of the storage space required for each of the N operation branches to run is the same, and will not be elaborated one by one.

[0116] Continuing with Figure 1 as an example, in addition to determining the storage space corresponding to the operation branch (such as the first operation branch) corresponding to operator a4 and the storage space corresponding to the operation branch (such as the second operation branch) corresponding to operator b4, the electronic device can also determine the storage space corresponding to the operation branch corresponding to operator c4 and the storage space corresponding to the operation branch corresponding to operator d4.

[0117] For example, the storage space corresponding to the operation branch corresponding to operator c4 is 9 units of storage space. The storage space corresponding to the operation branch corresponding to operator d4 is 10 units of storage space.

[0118] It can be understood that when each operator runs, the electronic device needs to obtain input data, and after the operation is completed, it also needs to store the obtained output data. Therefore, by using the data volume of the input data and / or output data of each operator as the storage space required during the operation of the operator, the accuracy of determining the storage space required during the operation of the operator can be higher, and further, the accuracy of determining the storage space required during the operation of each of the N running branches can also be higher.

[0119] S203: Run the N running branches in the order of the storage spaces corresponding to the N running branches from small to large, and obtain N running results corresponding to the N running branches one by one.

[0120] It can be understood that if the storage space corresponding to the first running branch is greater than the storage space corresponding to the second running branch, then run the second running branch first, and then run the first running branch. If the storage space corresponding to the first running branch is less than the storage space corresponding to the second running branch, then run the first running branch first, and then run the second running branch.

[0121] When running each of the N running branches, the embodiments of the present application do not limit the order of operation of the operators that do not have a dependency relationship during the operation of each running branch. For example, there is no dependency relationship between operator a1 and operator b1 in the running branch corresponding to operator b4 (for example, the second running branch). Therefore, operator a1 can be run first, and then operator b1, or operator b1 can be run first, and then operator a1.

[0122] Continuing with Figure 1 the model structure shown as an example, based on the content described in S202, it can be known that the storage space corresponding to the running branch corresponding to operator a4 is 4 units of storage space. The storage space corresponding to the running branch corresponding to operator b4 is 7 units of storage space. The storage space corresponding to the running branch corresponding to operator c4 is 9 units of storage space. The storage space corresponding to the running branch corresponding to operator d4 is 10 units of storage space.

[0123] Therefore, the running order of the four branches is: the running branch corresponding to operator a4, the running branch corresponding to operator b4, the running branch corresponding to operator c4, and the running branch corresponding to operator d4.

[0124] It can be understood that in the above manner, the electronic device only needs to determine the N storage spaces required during the operation of the N running branches once, and then it can determine the execution order of the N running branches. Then, during the operation of each running branch, run the corresponding running branch according to the execution order.

[0125] In some other embodiments, the electronic device may also update the storage space required for the remaining unexecuted running branches after each execution of a running branch, and then execute the running branch corresponding to the minimum storage space during the execution of the next running branch.

[0126] Continuing with the Figure 1 model structure shown as an example, after the electronic device determines the storage spaces corresponding to the running branches corresponding to operator a4, operator b4, operator c4, and operator d4 respectively, it determines that the storage space corresponding to the running branch corresponding to operator a4 is the smallest. Therefore, during the execution of the first running branch, the electronic device may execute the running branch corresponding to operator a4.

[0127] Then, the electronic device may update the storage spaces corresponding to the remaining unexecuted running branches (such as updating the running branches corresponding to operator b4, operator c4, and operator d4 respectively). Taking operator b4 as an example, since some operators in the running branch corresponding to operator b4 have been executed during the execution of the first running branch, when the running branch corresponding to operator b4 is executed at this time, there is no need for additional storage space to execute the operators that have already been executed (i.e., operator a1, operator a2, operator a3). Therefore, the storage space of the running branch corresponding to operator b4 is updated to 4 units of storage space. Similarly, the storage space corresponding to operator c4 is updated to 7 units of storage space. The storage space corresponding to operator d4 is updated to 9 units of storage space.

[0128] At this time, the storage space corresponding to operator b4 is the smallest, and the electronic device may execute operator b4 during the execution of the second running branch. And so on, until the running branches corresponding to operator c4 and operator d4 are executed.

[0129] It can be understood that there will be a corresponding execution result after each running branch is executed. Therefore, N running branches correspond to N execution results. The determination method for the execution result corresponding to each running branch will be described below.

[0130] Taking the determination of the execution result corresponding to the first running branch as an example, since the first running branch includes the operators that must be executed during the process from the first operator to the first target operator and have a direct or indirect dependency relationship with the first operator, the execution result corresponding to the first running branch is the execution result of the first target operator.

[0131] Taking the first target operator as operator a4 as an example, the electronic device may use the execution result of operator a4 as the execution result of the first running branch.

[0132] The determination process of the execution result of operator a4 will be described in detail below.

[0133] As Figure 1 shown, since there is a direct dependency between operator a1 and operator A, the electronic device can run operator a1 based on the operation result of operator A to obtain the operation result of operator a1. Since there is a direct dependency between operator a2 and operator a1, the electronic device can run operator a2 based on the operation result of operator a1 to obtain the operation result of operator a2. Similarly, the electronic device can run operator a3 based on the operation result of operator a2 to obtain the operation result of operator a3, and then run operator a4 based on the operation result of operator a3 to obtain the operation result of operator a4.

[0134] In some other embodiments, after the electronic device finishes executing S203 and obtains the N operation results corresponding to the N operation branches one by one, since the second operator directly depends on the N second intermediate operators, the electronic device can further run the second operator based on the N operation results to obtain the operation result of the second operator.

[0135] For example, the electronic device can be based on the operation result of the operation branch corresponding to operator a4, the operation result of the operation branch corresponding to operator b4, the operation result of the operation branch corresponding to operator c4, and the operation result of the operation branch corresponding to operator d4 on the basis of the Figure 1 shown model structure to run operator B and obtain the operation result of operator B. Thus, Figure 1 all the operators in the shown model structure have been executed.

[0136] In the method provided by the embodiments of the present application, since the first operation branch and the second operation branch include at least one same intermediate operator, if the storage space required by the second operation branch is greater than the storage space required by the first operation branch, it means that the probability that the intermediate operator of the second operation branch depends on the intermediate operator of the first operation branch is relatively high. If the first operation branch is run first and then the second operation branch is run at this time, since the same intermediate operator has been run when the first operation branch is run, the actual storage space occupied when the second operation branch is run is less, and thus the determination speed of the operation result of the second target operator can be improved.

[0137] In addition, in the method provided by this application, after the first execution branch finishes running, a part of the storage space corresponding to the first execution branch can be released. For other execution sequences, such as the execution sequence that first executes the second execution branch, the electronic device can only release a part of the storage space for the first time after the second execution branch finishes running. Since the storage space corresponding to the first execution branch is smaller than the storage space corresponding to the second execution branch, it indicates that the number of operators included in the first execution branch is probably less than the number of operators included in the second execution branch. Therefore, executing the first execution branch first can make the time when the execution branch finishes running earlier, and further make the time when a part of the storage space is released for the first time earlier, which is beneficial to the reuse of the storage space.

[0138] Although this application only describes the execution order of the first execution branch and the second execution branch by taking the size relationship of the storage space required when the first execution branch and the second execution branch in N execution branches run as an example, it should be understood that there may be other execution branches in the neural network model that meet the limiting conditions of the execution branches involved in the foregoing S201. And the electronic device can determine the storage space required when the execution branches that meet the above limiting conditions run according to the methods shown in the foregoing Figure 2 and Figure 3 to determine the execution order of the execution branches.

[0139] For example, among the N execution branches, there is also a fourth execution branch that runs from the first operator to the third target operator, and the fourth execution branch and the second execution branch include at least one same intermediate operator. Based on the model structure of the neural network model shown in Figure 1 , the third target operator can be, for example, operator c4. At this time, the fourth execution branch can include: operators a1 - a2, operators b1 - b3, operators c1 - c4.

[0140] For another example, among the N execution branches, there is also a fifth execution branch that runs from the first operator to the fourth target operator, and the fifth execution branch and the fourth execution branch include at least one same intermediate operator. Based on the model structure of the neural network model shown in Figure 1 , the fourth target operator can be, for example, operator d4. At this time, the fifth execution branch can include: operator a1, operators b1 - b2, operators c1 - c3, operators d1 - d4.

[0141] In some embodiments, the embodiments of this application also provide a readable storage medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device is caused to execute the running method of the neural network model described in the foregoing embodiments.

[0142] In some embodiments, the embodiments of the present application further provide a chip, which is used to implement the operation method of the neural network model described in the above embodiments.

[0143] In the present application, the chip may include one or more processing units. For example, it may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), a processing module or circuit of an artificial intelligence (AI) processor, or a field programmable gate array (FPGA). Among them, the AI processor includes a neural network processing unit (NPU), etc.

[0144] Taking the chip including an NPU as an example, the electronic device can run the neural network model in the present application based on the NPU and execute the operation method of the neural network model shown in the above Figure 2 and Figure 3 of the present application.

[0145] In some embodiments, the embodiments of the present application further provide an electronic device, which includes the above chip and one or more memories; one or more programs are stored in the one or more memories, and when the one or more programs are executed by the chip, the electronic device executes the operation method of the neural network model described in the above embodiments.

[0146] Figure 4 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown. As Figure 4 shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output (I / O) device 105, and a system control logic unit 106.

[0147] Among them: The processor 101 may include one or more processing units. For example, it may include a CPU, a GPU, a DSP, a MCU, a processing module or circuit of an AI processor, or an FPGA. Among them, the AI processor includes an NPU, etc.

[0148] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory 102 is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 can be used to store relevant instructions for the above method of running a neural network model, etc.

[0149] The non-volatile memory 103 can include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 can include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 can also be a removable storage medium, such as a secure digital (SD) memory card, etc.

[0150] Specifically, the system memory 102 and the non-volatile memory 103 can respectively include: a temporary copy and a permanent copy of the instruction 107. The instruction 107 can include: instructions that, when executed by at least one of the processors 101, cause the electronic device 100 to implement the method as Figure 2 and Figure 3 shown.

[0151] The communication interface 104 can include a transceiver for providing a wired or wireless communication interface for the electronic device 100, and then communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 can be integrated into other components of the electronic device 100. For example, the communication interface 104 can be integrated into the processor 101. In some embodiments, the electronic device 100 can communicate with other devices through the communication interface 104.

[0152] The input / output (I / O) device 105 can include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. The user can interact with the electronic device 100 through the input / output (I / O) device 105.

[0153] The system control logic unit 106 may include any suitable interface controllers to provide any suitable interfaces to other modules of the electronic device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide interfaces to the system memory 102 and the non-volatile memory 103.

[0154] In some embodiments, at least one of the processors 101 may be logically encapsulated with one or more controllers for the system control logic unit 106 to form a system in package (SiP). In some other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).

[0155] It can be understood that Figure 4 The structure of the illustrated electronic device 100 is only an example. In some other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0156] It can be understood that, as used herein, the term "module" may refer to or include an application-specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or grouped) that executes one or more software or firmware programs and / or a memory, combinational logic circuits, and / or other suitable hardware components that provide the described functionality, or may be part of these hardware components.

[0157] It can be understood that in various embodiments of the present application, the processor may be a microprocessor, a digital signal processor, a microcontroller, etc., and / or any combination thereof. According to another aspect, the processor may be a single-core processor, a multi-core processor, etc., and / or any combination thereof.

[0158] The various embodiments disclosed in the present application may be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application may be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0159] Program code can be applied to the input instructions to perform the various functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit, or a microprocessor.

[0160] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0161] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, a floppy disk, a compact disc, a CD-ROM, a magneto-optical disc, a ROM, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using electrical, optical, acoustic, or other forms of propagated signals via the Internet. Thus, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0162] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0163] It should be noted that each unit / module mentioned in the device embodiments of this application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by this application. In addition, to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed by this application. This does not mean that there are no other units / modules in the above device embodiments.

[0164] It should be noted that in the examples and descriptions of this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0165] Although this application has been illustrated and described by reference to certain embodiments thereof, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the scope of this application.

Claims

1. A method for operating a neural network model, characterized in that: Applied to electronic equipment, the method comprises: Acquire a neural network model to be run, wherein the neural network model includes a plurality of operators, and the plurality of operators include a first operator, a second operator, and K intermediate operators between the first operator and the second operator. M first intermediate operators among the K intermediate operators directly depend on the first operator, the second operator directly depends on N second intermediate operators among the K intermediate operators, there is no dependency relationship between the N second intermediate operators, and the M first intermediate operators are different from the N second intermediate operators. A first operation branch from the first operator to a first target operator among the N second intermediate operators and a second operation branch from the first operator to a second target operator among the N second intermediate operators include at least one identical intermediate operator, K is a positive integer greater than 3, and M and N are positive integers greater than 1; Determine the storage space required for each of the N running branches when the first operator runs to the N second intermediate operators respectively; Based on the order of the storage spaces corresponding to the N running branches from small to large, the N running branches are run to obtain N running results corresponding to the N running branches one by one.

2. The method according to claim 1, characterized in that The first running branch includes: operators that must be executed in the process of running from the first operator to the first target operator, and operators that have a direct dependency relationship or an indirect dependency relationship with the first operator; The second running branch includes: operators that must be executed in the process of running from the first operator to the second target operator, and operators that have a direct dependency relationship or an indirect dependency relationship with the first operator.

3. The method according to claim 1, characterized in that The storage space required for the third running branch among the N running branches when running includes: the storage space occupied by the third running branch in the memory and / or cache of the electronic device when running.

4. The method according to claim 3, characterized in that The storage space required when the third running branch is running includes: The sum of the amount of input data and / or output data of each operator in the third execution branch.

5. The method according to claim 1, characterized in that The method further comprises: After executing the N execution branches, the second operator is executed based on the N execution results to obtain an execution result of the second operator.

6. The method according to any one of claims 1 to 5, characterized in that: The multiple operators include one or more of the following operators: Arithmetic operators, convolution operators, pooling operators, normalization operators, fully connected operators, activation operators, loop operators, attention operators, tensor deformation operators, tensor concatenation and splitting operators, loss function operators, upsampling and interpolation operators.

7. The method according to any one of claims 1 to 5, characterized in that: The neural network model includes one or more of a deep learning model, a generative model, a machine learning model, and a reinforcement learning model.

8. A chip, characterized in that: The chip is used to implement the operating method of the neural network model described in any one of claims 1 to 7.

9. An electronic device, characterized in that: It comprises the chip as claimed in claim 8, and one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the chip, the electronic device executes the operating method of the neural network model as claimed in any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores instructions, which, when executed on an electronic device, enable the electronic device to execute the method for operating the neural network model described in any one of claims 1 to 7.

Citation Information

Cited By

  • Picture and information classification model generation method and device, equipment and storage medium

    CN115994318A

  • Picture, information classification model generation method and device, equipment and storage medium

    CN115994318B