Topology sorting method and device, equipment, AI platform and storage medium

By adjusting the lifecycle and execution order of nodes in the neural network model, the memory consumption problem caused by the uncertainty of the topology sorting result is solved, thereby reducing memory consumption and ensuring the determinism of the sorting result, and improving the efficiency and memory utilization of topology sorting.

CN121050874APending Publication Date: 2025-12-02HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410700877.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In existing technologies, neural network models have multiple possible sorting results during topological sorting, leading to uncertain memory usage and making it impossible to effectively reduce the memory consumption of neural networks.

Method used

By obtaining the first topological sorting result, the nodes that need to be reordered are identified, and their execution order is updated. This shortens the lifecycle of the nodes and reduces the number of nodes with overlapping lifecycles, thereby reducing memory usage and obtaining a definite topological sorting result.

Benefits of technology

By adjusting the lifecycle and execution order of nodes, the memory footprint of the neural network is reduced, the determinism and efficiency of topology sorting are improved, and the reasonable allocation of memory is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121050874A_ABST
    Figure CN121050874A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a topology sorting method and device, equipment, an AI platform and a storage medium, and relates to the technical field of neural network application. In the topological sorting process, after a first topological sorting result of the nodes in the neural network is obtained, first nodes needing delayed sorting are determined based on the first topological sorting result. And updating the first execution sequence of the first nodes to obtain a second execution sequence of the first nodes, and further updating the first topological sorting result according to the second execution sequence of the first nodes to obtain a second topological sorting result. According to the embodiment of the invention, the life cycle of the first node is shortened by determining the first node needing to be reordered and updating the first execution sequence of the first node. The memory occupation amount needing to be allocated is reduced, so that the effect of reducing the memory occupation amount of the neural network is achieved, and then a determined topological sorting result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neural network application technology, and more specifically to a topology sorting method, apparatus, device, artificial intelligence (AI) platform, and storage medium. Background Technology

[0002] With the rapid development of Artificial Intelligence (AI) technology, the scale of neural network models is constantly expanding, and the number of neurons in these models is also increasing dramatically. This leads to a growing demand for device memory resources during both the training and application phases of neural network models. When running neural network models on computing devices with relatively small memory capacities (such as mobile computing devices), this increased memory consumption can cause memory shortage problems. Therefore, minimizing the memory footprint of neural network models is desirable.

[0003] Currently, the main approach to reducing the memory footprint of neural networks is to perform topological sorting of the neurons involved in computation. However, in these techniques, multiple possible sorting results can occur during topological sorting, making the final result uncertain and thus failing to effectively reduce network memory usage. Summary of the Invention

[0004] This application provides a topology sorting method, apparatus, device, AI platform, and storage medium to obtain deterministic sorting results that can reduce network memory usage.

[0005] In a first aspect, embodiments of this application provide a topology sorting method. When sorting the nodes of a neural network topology, this method obtains a first topology sorting result, determines a first node that needs to be reordered based on the first topology sorting result, and updates the first execution order of the first node to obtain a second execution order of the first node. The lifetime of the first node in the second execution order is shorter than the lifetime of the first node in the first execution order. The first topology sorting result is then updated according to the second execution order of the first node to obtain a second topology sorting result.

[0006] Compared to related technologies where the sorting result is uncertain due to multiple possible sorting outcomes, this application's embodiment readjusts the order of the first node in the first topological sorting result. This is done so that the lifespan of the first node in the second execution order is shorter than its lifespan in the first execution order, thus shortening the lifespan of the first node. Consequently, the lifespan of the first node in the second topological sorting result is shorter than its lifespan in the first topological sorting result. By shortening the node lifespan, the number of nodes with overlapping lifespans is reduced, thereby reducing the amount of memory required. This results in the neural network using the second topological sorting result having less memory usage than the first topological sorting result, achieving a reduced memory footprint and ultimately obtaining a definite topological sorting result.

[0007] In one possible implementation, the first node is defined as a node in the first topological sorting result that satisfies the delayed sorting condition. The delayed sorting condition indicates the rules governing nodes that satisfy the delayed sorting criteria.

[0008] In this way, by determining the first node from the first topological sort result through the delayed sort condition, the node that needs to be delayed sorted can be located quickly and accurately, which improves the efficiency of topological sorting while ensuring the reliability of the second topological sort result.

[0009] In one possible implementation, the delayed sorting condition is used to indicate that the memory of a node is not reused. When determining the first node based on the first topological sorting result, the node whose memory is not reused is selected from the multiple nodes in the first topological sorting result, and the output node of the node whose memory is not reused is determined as the first node.

[0010] Because nodes in a neural network have different node types, some nodes reuse memory while others do not. For nodes that do not reuse memory, adjusting the execution order of their output nodes will not affect their memory lifespan. Therefore, in this possible implementation, nodes that do not reuse memory are selected from the first topological sorting result using a delayed sorting condition, and their output nodes are designated as the first node. This quickly locates the first node that needs delayed sorting, improving the efficiency of the topological sorting.

[0011] In one possible implementation, the delayed sorting condition is used to indicate the lifecycle threshold corresponding to a node. Here, lifecycle indicates the length of time between memory allocation and memory release. When determining the first node based on the first topology sorting result, the first node selected from the first topology sorting result whose lifecycle is greater than the corresponding lifecycle threshold is chosen.

[0012] Since longer node lifecycles increase the memory footprint of the neural network, this possible implementation identifies nodes with longer lifecycles in the first topological sorting result by using the lifecycle threshold indicated by the delayed sorting condition, thus delaying the sorting of these nodes. This reduces the memory footprint of the neural network.

[0013] In one possible implementation, the first node is determined based on the first topological sorting result. Specifically, this is achieved by selecting the first node from the second nodes whose lifetime is greater than the corresponding lifetime threshold. Optionally, the second node is a node in the first topological sorting result whose memory is reused.

[0014] If lifecycle identification were performed on every node in the first topology sorting result, the topology sorting would be extremely time-consuming when the number of nodes in the neural network is large. Furthermore, identifying the first node solely based on node type or lifecycle carries the risk of missed identifications, thus failing to effectively reduce the neural network's memory footprint. Therefore, in this possible implementation, memory is reused for second nodes selected from the first topology sorting result, and nodes with long lifecycles are chosen from these second nodes. This reduces the number of nodes requiring lifecycle comparison while improving the accuracy of the obtained first nodes, thereby ensuring the reliability of the second topology sorting result.

[0015] In one possible implementation, the lifecycle threshold for a node is the lifecycle of its output nodes. Alternatively, the lifecycle of a node can be the shortest lifecycle among its multiple output nodes.

[0016] Based on this possible implementation, when determining the first node based on the first topological sorting result, the node whose lifespan is greater than the lifespan of its output nodes can be determined as the first node. Alternatively, when a node has multiple output nodes, the node whose lifespan is greater than the shortest lifespan among its multiple output nodes can be determined as the first node.

[0017] In this way, by analyzing the lifecycle of each node's output, we can identify whether it is the first node and flexibly adjust the lifecycle threshold for each node. This improves the reliability of the lifecycle threshold for each node, thereby increasing the accuracy of obtaining the first node and ensuring the reliability of the second topological sorting result.

[0018] In one possible implementation, the first execution order of the first node is updated to obtain the second execution order of the first node. Specifically, based on the output node of the first node, the first execution order of the first node is updated to obtain the second execution order of the first node.

[0019] In this possible implementation, the first execution order of the first node is adjusted by adjusting the output node of the first node so that the interval between the first node and the output node in the second execution order is smaller than the interval between the first node and the output node in the first execution order, thereby shortening the life cycle of the first node.

[0020] In one possible implementation, the first execution order of the first node is updated to obtain the second execution order of the first node. Specifically, based on the input node of the first node, the first execution order of the first node is updated to obtain the second execution order of the first node.

[0021] In this possible implementation, the first execution order of the first node is adjusted by the input node of the first node, so that the interval between the first node and the input node in the second execution order is smaller than the interval between the first node and the input node in the first execution order, thereby shortening the life cycle of the input node.

[0022] In one possible implementation, the first execution order of the first node is updated based on its output node to obtain the second execution order. Specifically, this involves comparing the lifecycle of the first node with that of the output node. If the lifecycle of the first node is longer than that of the output node, the first execution order of the output node's preceding node is used as the second execution order of the first node. Optionally, the preceding node is the node whose first execution order is adjacent to the first execution order of the output node.

[0023] In this possible implementation, nodes with longer lifecycles are delayed and sorted before the corresponding output nodes, shortening their memory lifecycle without affecting the memory lifecycle of the input nodes, thereby reducing the memory usage of the neural network.

[0024] In one possible implementation, the process involves creating a node sequence for each node in the first topological sorting result before comparing the lifecycle of the first node with that of the output node. Each node sequence comprises multiple subsequences, with each node corresponding to one subsequence, and each subsequence containing at least one node. When the subsequence containing the first node contains a node, the lifecycle of the first node is compared with that of the output node.

[0025] Thus, when the node sequence containing the first node contains another node, by identifying whether the first node is a long-lived node, it is determined whether the execution order of the first node should be adjusted to shorten the lifespan of the identified long-lived node, thereby reducing the memory usage of the neural network.

[0026] In one possible implementation, the specific implementation is as follows: based on each node in the first topological sorting result, after creating a node sequence, when the subsequence where the first node is located contains multiple nodes, the first execution order of the preceding node of the output node is taken as the second execution order of the first node.

[0027] Thus, when the subsequence containing the first node contains multiple nodes, the first node in the subsequence is directly moved before the output node, simplifying the steps of determining the second execution order of the first node and improving the efficiency of topological sorting.

[0028] In one possible implementation, the first execution order of the first node is updated based on its input node to obtain a second execution order. Specifically, this involves comparing the lifecycle of the first node with that of the input node. If the lifecycle of the first node is shorter than that of the input node, the first execution order of the subsequent nodes of the input node is used as the second execution order of the first node. Optionally, the subsequent node is the node whose first execution order is adjacent to the first execution order of the input node.

[0029] In this possible implementation, nodes with long lifecycles are sorted later, shortening their memory lifecycle without affecting the memory lifecycle of input nodes, thereby reducing the memory usage of the neural network.

[0030] In one possible implementation, the first topological sorting result is obtained by configuring the execution order of each node in the directed acyclic graph (DAG) of the neural network to obtain the first topological sorting result.

[0031] In this way, by configuring the execution order of each node in the DAG of the neural network, the first topological sorting result is obtained, thus improving the efficiency of topological sorting.

[0032] In one possible implementation, the specific steps are as follows: After updating the first topological sorting result according to the second execution order of the first node to obtain a second topological sorting result, memory locations in computer-accessible memory are allocated to the nodes of the neural network according to the second topological sorting result. Data to be processed is input into the computer. The data to be processed includes at least one of text data, image data, or video data. The computer performs processing operations on the data to be processed and outputs the processing result. The processing result includes at least one of recognition result or generated content.

[0033] In this way, by allocating memory locations based on the second topological sorting result, the memory usage of the neural network can be reduced.

[0034] In one possible implementation, allocating computer-accessible memory locations to nodes of the neural network based on the second topology sorting result is specifically implemented as follows: determining the lifecycle of nodes in the neural network based on the second topology sorting result; and allocating memory locations to nodes in the neural network based on the lifecycle of nodes in the DAG of the neural network.

[0035] In this possible implementation, memory locations are allocated to nodes with non-overlapping lifecycles based on the lifecycles of nodes in the DAG of the neural network. This reduces the memory footprint of the neural network.

[0036] Secondly, embodiments of this application provide a topology sorting device, which includes a sorting module, a node determination module, an order correction module, and a topology determination module.

[0037] The sorting module is used to obtain the first topological sorting result; the first topological sorting result is used to indicate the first execution order of nodes in the neural network.

[0038] The node determination module is used to determine the first node based on the first topological sorting result. The first node is a node that has been reordered in the neural network.

[0039] The sequence correction module is used to update the first execution order of the first node to obtain the second execution order of the first node. Optionally, the lifespan of the first node under the second execution order is shorter than the lifespan of the first node under the first execution order.

[0040] The topology determination module is used to update the first topology sorting result according to the first execution order of the first node to obtain the second topology sorting result.

[0041] Thirdly, embodiments of this application provide an AI platform for providing users with services for building and running neural networks. This AI platform is used to obtain the topological sorting result of nodes in the neural network through the first aspect or any possible implementation thereof. Based on the topological sorting result of the nodes in the neural network, memory locations in computer-accessible memory are allocated to the nodes of the neural network. Data to be processed is input into the computer. The computer performs processing operations on the data to be processed and outputs the processing result of the data to be processed.

[0042] Optionally, the data to be processed includes at least one of text data, image data, or video data. The processing operation is used to identify or generate at least one of these.

[0043] Fourthly, embodiments of this application provide a computing device including a processor coupled to a memory. The processor executes instructions stored in the memory to cause the computing device to perform the methods described in the first aspect or any possible implementation thereof.

[0044] Fifthly, embodiments of this application provide a computer program product containing instructions that, when executed by a computing device, cause the computing device to perform the method described in the first aspect or any possible implementation thereof.

[0045] In a sixth aspect, embodiments of this application provide a computer-readable storage medium including computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the instructions in the computer program stored in the computer-readable storage medium to perform the method in the first aspect or any possible implementation of the first aspect.

[0046] The technical effects of any implementation of aspects two through six can be found in aspect one or any possible implementation of aspect one. Further details are omitted here.

[0047] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0048] Figure 1 This is a diagram illustrating topological sorting.

[0049] Figure 2 This is a schematic diagram of the system architecture of the topology sorting method provided in the embodiments of this application;

[0050] Figure 3 A flowchart illustrating the topology sorting method provided in an embodiment of this application;

[0051] Figure 4 A schematic diagram illustrating the construction of a directed acyclic graph (DAG) of nodes in a neural network provided in an embodiment of this application;

[0052] Figure 5 A schematic diagram illustrating the acquisition of the first topological sorting result based on DAG, provided in an embodiment of this application;

[0053] Figure 6 A flowchart illustrating the determination of the second execution order of the first node, provided for an embodiment of this application;

[0054] Figure 7 This is a schematic diagram of the node sequence provided in an embodiment of this application;

[0055] Figure 8AA schematic diagram illustrating the adjustment of the execution order of a first node, provided for an embodiment of this application;

[0056] Figure 8B A schematic diagram illustrating another adjustment to the execution order of the first node provided in an embodiment of this application;

[0057] Figure 9 This is a schematic diagram of the node sequence after rearrangement provided in an embodiment of this application;

[0058] Figure 10 A flowchart illustrating the process of determining the second topological sorting result provided in an embodiment of this application;

[0059] Figure 11 A schematic diagram of memory location allocation for a node provided in an embodiment of this application;

[0060] Figure 12 This is a schematic diagram of the topology sorting device provided in an embodiment of this application;

[0061] Figure 13 A schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0062] First, the terms used in the embodiments of this application will be explained.

[0063] AI platform: A platform that provides AI developers and users with a convenient AI development environment and convenient development tools.

[0064] The AI ​​platform has a built-in deep learning framework, which allows users to build, train, and infer neural network models.

[0065] Deep learning frameworks: Platforms used for developing and running artificial intelligence algorithms. Examples include TensorFlow, Caffe, and PyTorch. These frameworks provide users with various functionalities such as data input, neural network model creation, model training, model execution, and hardware drivers.

[0066] Neural network model: A model built using neural networks to process tasks. Tasks handled by this model may include image recognition, image classification, text-to-image conversion, etc.

[0067] In a neural network model, multiple neurons perform calculations on the input data to achieve task processing. As the size of the neural network model increases, the number of neurons also increases, leading to a greater consumption of memory resources. To reduce the memory footprint of neural network models, current methods primarily involve network pruning, memory reuse, or controlling the execution order to reduce memory size.

[0068] Network pruning refers to the technique of reducing the memory size of a neural network model by decreasing the number of neurons in the model.

[0069] Memory reuse refers to the technique of allowing different neurons in a neural network model to share the same memory, thereby reducing memory overhead by implementing time-sharing memory reuse.

[0070] Here, different neurons can refer to neurons whose life cycles do not overlap.

[0071] The lifecycle refers to the length of time from memory allocation to memory release. It includes three stages: memory allocation, memory usage, and memory release. The lifecycle can also be called the memory lifecycle, the neuron lifecycle, or the node lifecycle. In a neural network model, a node can refer to a neuron, a variable, or a computation node.

[0072] For example, neuron A and neuron B have non-overlapping lifecycles. After a certain memory location (memory 1) is allocated to neuron A, and neuron A completes its data processing, that memory location (memory 1) is released and allocated to neuron B. In this way, neuron A and neuron B share the same memory location (memory 1).

[0073] Memory reuse can reduce memory usage without affecting the performance of the neural network, making the memory usage of the neural network model approach the theoretical minimum memory usage.

[0074] Memory usage refers to the amount of memory space required by a neural network when performing data processing. Controlling the execution order, also known as topology sorting optimization, refers to techniques that reduce the theoretical minimum memory usage of a neural network model by modifying the topological sorting of its neurons.

[0075] The execution order can refer to the execution timing or sequence of neurons in a neural network.

[0076] Topological sorting refers to the process of ordering the nodes in a Directed Acyclic Graph (DAG). Topological sorting is used to determine the execution order of nodes in a DAG. In this embodiment, topological sorting is used to order nodes in a neural network model. Nodes in the neural network model can refer to neurons, variables, or computational nodes.

[0077] Currently, topological sorting can be achieved through Depth First Search (DFS), Breadth First Search (BFS), and Reverse Depth First Search (RDFS).

[0078] For example, taking DFS as an example, such as Figure 1 As shown, Figure 1 This is a schematic diagram of topological sorting. Figure 1 Figure (b) is a schematic diagram of a Directed Acyclic Graph (DAG), which includes six nodes: A, B, C, D, E, and F. The output of node A serves as the input to nodes B and C, respectively. The output of node B serves as the input to node F. Node C is subsequently connected to nodes D and E, and the output of node E serves as the input to node F. Since... Figure 1 The DAG graph shown in Figure (b) has two branches, for Figure 1 When performing a depth-first search on graph (b), if the branch containing node B is sorted first, the result will be as follows: Figure 1 The topological sorting result shown in Figure (c) is {A,B,C,D,E,F}. Figure 1 As shown in diagram (c), the execution order of nodes A, B, C, D, E, and F is 1, 2, 3, 4, 5, and 6, respectively. If the branch containing node C is sorted first, the result will be as follows: Figure 1 The topological sorting result shown in Figure (d) is {A,C,D,E,B,F}. Figure 1 As shown in Figure (d), the execution order of nodes A, B, C, D, E and F is 1, 5, 2, 3, 4 and 6 respectively.

[0079] from Figure 1 It can be seen that when a DAG graph has branches, the current topology ranking method will produce at least two possible topology ranking results, making the result uncertain. For example, neural networks with the same DAG may have different topology ranking results, or the same neural network may produce different topology ranking results each time it is topologically ranked at different time periods. That is, for Figure 1 When performing topological sorting on the DAG graph shown in Figure (b), the following may occur: Figure 1 The topological sorting result shown in Figure (c) may also be as follows: Figure 1 The topological sorting result of graph (d) in the diagram. And... Figure 1 The topological sorting results shown in Figure (c) and Figure 1 The topological sorting result of graph (d) in the diagram appears randomly.

[0080] Different topological sorting results will lead to different memory reuse outcomes. For example... Figure 1As shown in Figure (c), sorting the branch containing node B first results in a lifetime of node B of [2,6]. The memory occupied by node B will not be released until node F is executed, making the lifetime of node B relatively long. This lifetime overlaps with the lifetimes of nodes C and D, thus preventing memory reuse with nodes C and D. However, when sorting the branch containing node C first, as shown in Figure (c), the lifetime of node B is [2,6]. Figure 1 As shown in Figure (d), the lifecycle of node B is [5,6]. Since the lifecycles of node B, C, and D do not overlap, node B can reuse memory with nodes C and D. However, because the lifecycle of node B is [5,6], the memory occupied by node A will only be released after node B executes, causing the lifecycle of node A to become [1,5]. It is evident that in related technologies, the topology sorting method leads to uncertainty in the topology sorting result, resulting in uncertainty in the lifecycles of nodes in the topology sorting result, and thus uncertainty in the amount of memory to be allocated, ultimately leading to an uncertain result in the size of the reused memory.

[0081] Based on this, to obtain a deterministic sorting result that reduces the memory consumption of the neural network, this application provides a topology sorting method. In the topology sorting process, after obtaining the first topology sorting result of the nodes in the neural network, this method determines the first node that needs to be delayed in sorting based on the lifecycle of the nodes in the first topology sorting result. Then, based on the output node corresponding to the first node, the original execution order of the first node is updated to obtain the execution order of the first node, such that the lifecycle of the first node in the second execution order is shorter than the lifecycle of the first node in the original execution order. The first topology sorting result is then updated according to the execution order of the first node to obtain a second topology sorting result. This application determines the first node that needs to be delayed in sorting by its lifecycle and, based on the output node corresponding to the first node, readjusts the sorting of the first nodes in the first topology sorting result, shortening the lifecycle of the first nodes, so that the lifecycle of the first node in the second topology sorting result is shorter than the lifecycle of the first node in the first topology sorting result. In this way, by shortening the lifecycle of nodes and reducing the number of nodes with overlapping lifecycles, the amount of memory that needs to be allocated is reduced. As a result, the memory usage of the neural network with the second topology sorting result is less than that with the first topology sorting result. This achieves the effect of reducing the memory usage of the neural network and thus obtaining a definite topology sorting result.

[0082] The topology sorting method provided in this application can perform topology sorting of the execution order of neurons in a neural network model during the training phase, and reuse memory based on the topology sorting result, thereby reducing the memory occupied by the neural network model during training. Alternatively, the topology sorting method provided in this application can perform topology sorting of the execution order of neurons in a trained neural network model during the application phase, and reuse memory based on the topology sorting result, thereby reducing the memory occupied by the neural network model during inference.

[0083] The topology sorting method provided in this application can be applied to AI platforms or deep learning frameworks. This application does not limit its application in this regard.

[0084] For example, taking the application of the topology sorting method to an AI platform, the system architecture of the topology sorting method is provided. Figure 2 As shown, Figure 2 This is a schematic diagram of the system architecture of the topology sorting method provided in this application embodiment. The system architecture of the topology sorting method shown includes a client 20 and a server 10. An AI platform 101 is deployed in the server 10. The server 10 provides neural network model development and operation services to the client 20 through the deployed AI platform 101.

[0085] Client 20 and server 10 establish a connection via network 30. Network 30 provides a communication link between client 20 and server 10. Network 30 can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 20 may require encoding, transcoding, compression, or other processing before being published to server 10.

[0086] In one possible implementation, client 20 can be a browser, an application (APP), a web application such as an H5 application or a lightweight application, or a cloud application. Client 20 is obtained through software development based on the corresponding services provided by server 10. Client 20 can be deployed on a terminal device and depends on the terminal device or certain APPs on the device to run. The terminal device can have a display screen and support information browsing; for example, the terminal device can be a mobile phone, tablet computer, or personal computer. Various other types of applications can typically be configured on the terminal device, such as human-computer interaction applications, model training applications, text data applications, web browser applications, shopping applications, search applications, cloud desktop applications, and cloud computer applications.

[0087] In one possible implementation, server 10 includes multiple instances. An instance can refer to a virtual machine, container, bare metal server, physical server, or similar entity that contains computing components such as CPU, memory, operating system, network, and disk.

[0088] In one possible implementation, server 10 can be a cloud server. When server 10 is a cloud server, it can include a cloud data center and a cloud service platform. The cloud data center deploys an AI platform 101.

[0089] In one example, AI platform 101 can be deployed independently in an instance within a cloud data center. Alternatively, AI platform 101 can be deployed in a distributed manner across multiple instances within a cloud data center.

[0090] AI platform 101 is abstracted into an AI cloud service by the cloud service provider and provided to users. After the user purchases the cloud service through client 20 on the cloud service platform (pre-payment is possible and settlement is made based on the final resource usage), the cloud environment uses AI platform 101 deployed in the cloud data center to provide the user with neural network model development and operation services.

[0091] In another possible implementation, server 10 can be a local server. When server 10 is a local server, it includes servers that provide services. For example, a server that provides data forwarding services to client 20; a server that provides neural network model building services to client 20; or a server that provides neural network model training services to client 20. It should be noted that server 10 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain.

[0092] In one example, AI platform 101 can be deployed independently on a server of server 10. Alternatively, AI platform 101 can be deployed in a distributed manner on multiple servers of server 10.

[0093] When using the neural network model development and operation services provided by AI Platform 101, users can specify the tasks to be performed by the neural network model and upload training datasets to AI Platform 101 through the application program interface (API) or graphical user interface (GUI) on client 20. AI Platform 101 receives the user's task information and training dataset, performs data preprocessing, and trains the neural network model. AI Platform 101 returns status information of the neural network model's training process to the user through the API or GUI. The trained neural network model can be downloaded from client 20 or used online to complete specific tasks, such as image recognition and image generation.

[0094] In the first possible implementation, in practical applications, the system architecture may include a server 10 and multiple clients 20. The multiple clients 20 establish communication connections through the server 10. The server 10 provides development and execution services for neural network models to the multiple clients 20.

[0095] Alternatively, multiple clients 20 can each act as a sender or a receiver. Communication between the multiple clients 20 is achieved through the server 10. Users can interact with the server 10 through the client 20 to receive data sent by other clients 20, or to send data to other clients 20, etc.

[0096] For example, during the neural network model training phase, client A determines the task the neural network model needs to complete and uploads the training dataset to server 10. Server 10 receives the user's task information and training dataset through the deployed AI platform 101, performs data preprocessing and neural network model training, and forms a trained neural network model. Server 10 then sends the trained neural network model to client B.

[0097] In a second possible implementation, the client 20 may have functions similar to those of the server 10, thereby executing the topology sorting method provided in the embodiments of this application to determine the topology sorting result of the nodes in the neural network model running in the client 20.

[0098] For example, client B deploys an AI platform 101. Client A requests the development and operation services of a neural network model from client B through server 10. Client B executes the topology sorting method provided in this application embodiment through AI platform 101 to determine the topology sorting result of the nodes in the neural network model running in client 20, thereby obtaining the trained neural network model or processing result. Client B returns the trained neural network model or processing result to client A through server 10.

[0099] For example, client B deploys an AI platform 101. Client A requests development and operation services for a neural network model from client B. Client B executes the topology sorting method provided in this application embodiment through AI platform 101 to determine the topology sorting result of the nodes in the neural network model running in client 20, thereby obtaining the trained neural network model or processing result. Client B returns the trained neural network model or processing result to client A.

[0100] For example, client A deploys an AI platform 101. Client A executes the topology sorting method provided in this application embodiment through the AI ​​platform 101 to determine the topology sorting result of the nodes in the neural network model running on the client, thereby obtaining the trained neural network model or processing result. Client A returns the trained neural network model or processing result to client B.

[0101] For example, taking server 10 as an example, where AI platform 101 is deployed, such as... Figure 2 As shown, the AI ​​platform 101 includes a network model layer 1011, a deep learning framework layer 1012, and a hardware layer 1013.

[0102] The network model layer 1011 is used to provide a neural network model and to obtain the neural network model selected or built by the user.

[0103] The deep learning framework layer 1012 is used to compile user-selected or built neural network models and run the compiled neural network models.

[0104] Hardware layer 1013 provides the hardware resources needed for the neural network model to run, such as an NPU, CPU, or GPU.

[0105] In one possible implementation, when the deep learning framework layer 1012 runs the compiled neural network model, it executes the topology sorting method provided in this application embodiment to determine the topology sorting result of the nodes in the compiled neural network model. Memory is allocated according to the topology sorting result of the nodes in the compiled neural network model. A running task for the neural network model is generated, and hardware resource allocation is performed based on the running task to determine the corresponding hardware for executing the running task. The running task is then executed through the hardware.

[0106] In one example, such as Figure 2 As shown, the deep learning framework layer 1012 includes an adaptation unit 10121, a model generation unit 10122, and a model execution unit 10123.

[0107] The adaptation unit 10121 is used to compile the neural network model selected or built by the user. It converts the user-selected or built neural network model into code that can be executed by the hardware in the hardware layer 1013.

[0108] The model generation unit 10122 is used to execute the topology sorting method provided in the embodiments of this application to determine the topology sorting result of the nodes in the compiled neural network model. It allocates memory according to the topology sorting result of the nodes in the compiled neural network model and generates the running task of the neural network model.

[0109] The model execution unit 10123 is used to allocate hardware resources based on the running task and determine the corresponding hardware for executing the running task. The running task is then executed through the hardware.

[0110] It should be noted that, Figure 2 The component divisions and naming in the provided system architecture are illustrative. In practical applications, they can be more complex than... Figure 2 More or fewer components.

[0111] based on Figure 2 The provided system architecture, in this application embodiment, offers a topology sorting method. This topology sorting method can be... Figure 2 The AI ​​platform 101 in the middle is executed. Or by Figure 2 The deep learning framework is executed in layer 1012. Figure 2 Taking the AI ​​platform 101 in the middle as an example, the topology sorting method provided in the embodiments of this application will be introduced.

[0112] like Figure 3 As shown, Figure 3 This is a schematic flowchart illustrating the topology sorting method provided in an embodiment of this application. The topology sorting method shown includes steps S310 to S340.

[0113] S310, obtain the first topological sort result.

[0114] The first topological sorting result is used to indicate the first execution order of each node in the DAG of the neural network. It includes each node in the DAG of the neural network and the first execution order of each node. The execution order indicates the chronological order in which the nodes are executed. A lower execution order indicates that the node is executed first, while a higher execution order indicates that the node is executed later.

[0115] The Directed Acyclic Graph (DAG) of a neural network, as shown above, can include multiple nodes, each corresponding to a neuron, variable, or computation node within the neural network. The specific process for generating a DAG of a neural network can be found below. Figure 4 The embodiments provided are not described in detail here.

[0116] For example, taking the neurons in the neural network corresponding to each node as an example, such as Figure 1 As shown, the connection method of neurons in a neural network is as follows: Figure 1 As shown in Figure (a), the output of neuron A is input into neurons B and C, respectively. The output of neuron B is input into neuron F. The output of neuron C is input into neuron D. The output of neuron D is input into neuron E. The output of neuron E is input into neuron F. Neuron F takes the superposition of the outputs of neuron B and neuron E as its input and processes it. Figure 1The schematic diagram of the DAG of the neural network shown in Figure (a) is as follows: Figure 1 Figure (b) is based on... Figure 1 The neural network shown in Figure (a) produces the following first topological sorting result: Figure 1 Figure (c) or as shown Figure 1 Figure (d) in the middle.

[0117] In one possible implementation, the execution order of each node in the DAG of the neural network can be configured using a topological sorting algorithm to obtain the first topological sorting result.

[0118] For example, at least one of the topological sorting algorithms, such as depth-first search, breadth-first search, or reverse depth-first search, can be used to perform topological sorting on the nodes in the DAG of a neural network. This yields the execution order of each node, and the nodes are then sorted sequentially according to the chronological order of their execution to obtain the first topological sort result. Specifically, the execution process of these topological sorting algorithms can be found in existing technologies and will not be elaborated upon here.

[0119] S320, determine the first node based on the first topological sorting result.

[0120] The first node comprises a reordered node from among multiple nodes in the neural network. This reordering involves either moving the execution order of nodes backward or forward, essentially modifying the execution time of the nodes.

[0121] It should be noted that if the first node exists in the first topological sort result, execute steps S330 to S340 below to obtain the second topological sort result. If the first node does not exist in the first topological sort result, it means that the first topological sort result has the minimum memory usage and does not require adjusting the execution order of the nodes; therefore, the first topological sort result is determined as the second topological sort result.

[0122] In a first possible implementation, the first node can be determined from multiple nodes in the first topological sorting result based on the node's lifecycle. For example, the node with the longest lifecycle is determined as the first node. Here, the node's lifecycle characterizes the length of time from when the node's memory is allocated to when it is released. That is, the length of time between when the node begins execution and when it ends execution. In this embodiment, the node's lifecycle can be represented by a range, the start of which is the execution order of the nodes, and the end of which is the execution order of the node's output nodes. For example, as... Figure 1 As shown in Figure (c), node B occupies memory when executing sequence 2, and the memory of node B is released when node F executes. Since the execution sequence of node F is 6, the lifetime of node B is [2,6].

[0123] Understandably, when a node has a long lifespan, there's a high probability that its lifespan overlaps with that of other nodes in the neural network. This prevents the node's memory from being reused with other nodes, requiring memory allocation for each node with overlapping lifespans, resulting in high memory usage for the neural network. Therefore, by identifying the node with the longest lifespan as the first node and adjusting its execution order, the first topological sorting result is corrected, shortening the first node's lifespan. This reduces the number of nodes with overlapping lifespans, leading to a topological sorting result with minimal memory usage.

[0124] In one example, the lifecycle of a node can be compared with the corresponding lifecycle threshold, and the node whose lifecycle is greater than the corresponding lifecycle threshold can be identified as the first node.

[0125] The lifecycle threshold can be preset or set based on the lifecycle of the input node or the lifecycle of the output node of each node.

[0126] For example, consider setting a lifetime threshold for a node based on the lifetime of its input nodes. For each node in the neural network, the lifetime of its input nodes under the first topological sorting result is used as the lifetime threshold for that node. Alternatively, when a node has multiple input nodes, the shortest lifetime among the lifetimes of the multiple input nodes can be used as the lifetime threshold for that node.

[0127] For example, consider setting a lifespan threshold for a node based on the lifespan of its output nodes. For each node in the neural network, the lifespan of its output nodes under the first topology ranking result is used as the lifespan threshold for that node. Alternatively, if a node has multiple output nodes, the shortest lifespan among the multiple output nodes can be used as the lifespan threshold for that node. In yet another example, the lifespan of each node in the first topology ranking result can be obtained, and the node with the longest lifespan among all nodes in the first topology ranking result can be determined as the first node.

[0128] In the second possible implementation, a probability prediction can be performed on each node to obtain the predicted probability that each node is a reordered node. Nodes with predicted probabilities greater than a probability threshold are identified as the first node.

[0129] Optionally, a probability prediction can be made for each node using a prediction model. This prediction model includes, but is not limited to, machine learning models and neural network models.

[0130] S330, update the first execution order of the first node to obtain the second execution order of the first node.

[0131] The lifespan of the first node in the second execution order is shorter than that of the first node in the first execution order.

[0132] In the first possible implementation, when the reordering shifts the execution order of nodes backward, the first execution order of the first node can be updated based on its output node to obtain the second execution order. Correspondingly, the output lifetime of the first node in the second execution order is shorter than that in the first execution order. Here, the output lifetime can refer to the range starting from the node's execution order and ending with the execution order of its output nodes.

[0133] In the second possible implementation, when the reordering moves the execution order of nodes forward, the first execution order of the first node can be updated based on its input nodes to obtain the second execution order. Accordingly, the input lifetime of the first node in the second execution order is shorter than that in the first execution order. The input lifetime can refer to the range starting from the execution order of the node's input nodes and ending with the node's execution order. Understandably, the input lifetime of a node is the same as the output lifetime of its input nodes. It should be noted that, in the following text, the lifetime of a node generally refers to its output lifetime.

[0134] In this context, the output node is the node in the neural network that takes the output of the first node as its input; that is, the node that receives the output of the first node. Figure 1 Taking diagram (b) as an example, when node B is the first node, the output of node B is sent to node F as the input of node F. Therefore, node F is the output node corresponding to node B. The input node is the node in the neural network whose output is used as the input of the first node. Figure 1 Taking Figure (b) as an example, when node B is the first node, the output of node A will be sent to node B as the input of node B. Therefore, node A is the input node corresponding to node B.

[0135] In one possible implementation, taking updating the first execution order of the first node based on the output node of the first node as an example, the first execution order of the first node can be updated based on the first execution order of the output node of the first node to obtain the second execution order of the first node.

[0136] In one example, when the first node has an output node, the second execution order of the first node is obtained based on the first execution order of the output node of the first node.

[0137] For example, the execution order preceding the first execution order of the output node can be determined as the second execution order of the first node, thereby moving the execution order of the first node before the output node. Figure 1 Taking Figure (c) as an example, when node B is the first node, node F is the output node of node B. The execution order of node F is 6, so the execution order 5 is taken as the second execution order of node B.

[0138] In another example, when the first node has multiple output nodes, the target output node with the highest execution order is determined from among the multiple output nodes. The second execution order of the first node is then updated based on the first execution order of this target output node.

[0139] In one possible implementation, taking updating the first execution order of the first node based on the input nodes of the first node as an example, the first execution order of the first node can be updated based on the first execution order of the input nodes of the first node to obtain the second execution order of the first node.

[0140] In one example, when the first node has an input node, the second execution order of the first node is obtained based on the first execution order of the input node of the first node.

[0141] For example, the execution order of the first node can be determined as the second execution order of the first node, thereby moving the execution order of the first node after the input node. Figure 1 Taking the diagram in (c) as an example, when node C is the first node, node A is the input node of node C. The execution order of node A is 1, so the execution order 2 is taken as the second execution order of node C.

[0142] In another example, when the first node has multiple input nodes, the target input node with the latest execution order is determined from among the multiple input nodes. The second execution order of the first node is then updated based on the first execution order of this target input node.

[0143] S340, update the first topology sorting result according to the second execution order of the first node to obtain the second topology sorting result.

[0144] The second topological sorting result is used to indicate the second execution order of the nodes in the neural network.

[0145] Since the lifespan of the first node in the second execution order is shorter than that in the first execution order, meaning the lifespan of the first node in the second topological sorting result is shorter than that in the first topological sorting result, the number of nodes with overlapping lifespans in the second topological sorting result is reduced. This increases the number of nodes that can reuse memory, thereby reducing the memory allocation size. Consequently, the memory usage of the second topological sorting result is less than that of the first topological sorting result. Thus, by adjusting the execution order of the first node, the memory usage of the topological sorting result is reduced.

[0146] In one possible implementation, updating the first topological sorting result according to the second execution order of the first node includes: retaining the first execution order of the preceding nodes of the first node in the first topological sorting result, updating the first execution order of the first node in the first topological sorting result to the second execution order of the first node, and adjusting the first execution order of the subsequent nodes of the first node in the first topological sorting result.

[0147] In this context, the preceding node is the node whose first execution order in the first topological sorting result precedes that of the first node; that is, the node whose first execution order in the first topological sorting result is less than that of the first node. The subsequent node is the node whose first execution order in the first topological sorting result follows that of the first node; that is, the node whose first execution order in the first topological sorting result is greater than that of the first node. For example, taking the first topological sorting result as... Figure 1 Taking the topological sorting result shown in Figure (c) as an example, when node B is the first node, then node A is the predecessor node of the first node. Nodes C, D, E, and F are the successors of the first node.

[0148] In one example, adjusting the first execution order of subsequent nodes of the first node in the first topological sorting result includes: reordering the first execution order of subsequent nodes of the first node based on the first execution order of the first node in the first topological sorting result to obtain a second execution order of the subsequent nodes. For example, taking the first topological sorting result as... Figure 1 Taking the topological sorting result shown in Figure (c) as an example, when node B is the first node, after setting the second execution order of node B to 5, the branch containing node C is executed first. The second execution orders of nodes C, D, and E are then set to 2, 3, and 4 respectively, resulting in the following... Figure 1 The second topological sorting result is shown in Figure (d).

[0149] based on Figure 3In the provided embodiment, during the topology sorting process, after obtaining the first topology sorting result of the nodes in the neural network, the first node that needs to be reordered is determined from the first topology sorting result. The first execution order of the first node is then updated to obtain a second execution order, so that the lifespan of the first node in the second execution order is shorter than that in the first execution order. The first topology sorting result is then updated according to the execution order of the first node to obtain the second topology sorting result. This embodiment of the application, by determining the first node that needs to be reordered and readjusting the order of the first nodes in the first topology sorting result, shortens the lifespan of the first node, making the lifespan of the first node in the second topology sorting result shorter than that in the first topology sorting result. Thus, by shortening the lifespan of nodes, the number of nodes with overlapping lifespans is reduced, thereby reducing the amount of memory required. This results in the neural network using the second topology sorting result having less memory usage than using the first topology sorting result, achieving the effect of reducing the memory usage of the neural network and obtaining a determined topology sorting result.

[0150] Next, we will introduce how to obtain the DAG of nodes in a neural network.

[0151] In one possible implementation, the directed acyclic graph of nodes in a neural network can be constructed using the graph of the neural network.

[0152] In neural networks, the graph is used to describe the connections between nodes and the data transmission relationships within the network. For example... Figure 4 As shown in Figure (a), Figure 4 Figure (a) is a schematic diagram of a neural network. Each point represents a node in the neural network. Figure 4 As shown in Figure (a), nodes A, B, and C compute in parallel. The outputs of nodes A, B, and C are input to nodes D and E. The output of node D is input to nodes F and G. The output of node E is input to node H. The output of node F is input to node I, and the output of node G is input to node J. The output of node H serves as the input to node K. The outputs of nodes I and neuron J are input to node L. The output of node K is input to node M.

[0153] In one implementation, the designed deep learning model file can be saved in formats such as tflite, pb, or onnx using mainstream deep learning frameworks such as tensorflow and pytorch, thus obtaining the graph of the neural network model.

[0154] In one implementation, after obtaining the graph of the neural network, a DAG of the neural network is created through the data transfer relationships between nodes in the neural network.

[0155] In another embodiment, after obtaining the graph of the neural network, an initial DAG of the neural network is created through the data transmission relationship between nodes in the neural network. The starting node and the ending node are then expanded in the initial DAG of the neural network to obtain the DAG of the neural network.

[0156] For example, such as Figure 4 As shown, based on Figure 4 Figure (a) shows the data transfer relationships between nodes in a neural network. The initial DAG of the neural network is created as follows: Figure 4 As shown in Figure (b), in the initial DAG of the neural network, a starting node X and a ending node Y are added, and the edges between the starting node X and nodes A, B, and C are expanded, and the edges between nodes M and L and the ending node Y are expanded, resulting in the following: Figure 4 The DAG of the neural network shown in Figure (c).

[0157] In this embodiment of the application, after obtaining the DAG of the neural network, the execution order of each node in the DAG of the neural network is configured to obtain the first topological sorting result.

[0158] In the first possible implementation, the execution order of each node in the DAG of the neural network can be configured with reference to the above S310.

[0159] In the second possible implementation, during topological sorting, the execution order of each node in the DAG of the neural network is set in ascending order of memory usage.

[0160] In one example, using depth-first search, setting the execution order of each node in the DAG of a neural network in ascending order of memory usage includes: creating an empty sequence, sequentially searching the nodes in the DAG, and writing them into the empty sequence. When branches exist, the search order of the branches is determined in ascending order of memory usage. Following the branch search order, nodes in each branch are written into the empty sequence one by one until all nodes in the entire DAG of the neural network are connected to the empty sequence, resulting in a sorted sequence indicating the execution time order of each node in the DAG of the neural network. This sorted sequence is determined as the first topological sort result.

[0161] For example, taking depth-first search as an example, such as Figure 5 As shown, the DAG of the neural network is as follows: Figure 5As shown. Create a sequence S1 = {}, which is initially empty. When an executable node A is found, its memory usage is X1. Add node A to sequence S1, making S1 = {A}. When the next executable node B is found, its memory usage is X2. Add node A to sequence S1, making S1 = {A,B}. When the next executable nodes C and D are found, node C has a memory usage of X3, and node D has a memory usage of Y1. If X3 is greater than Y1, add node D to sequence S1, making S1 = {A,B,D}. Then perform a depth-first search on the branch containing node D. After adding node E from the branch containing node D to sequence S1, perform a depth-first search on the branch containing node C, continuing until all nodes in the DAG of the neural network are added to sequence S1. The final sorted sequence is {A,B,D,E,C,F,G}.

[0162] In this embodiment of the application, after obtaining the first topological sorting result, the first node is determined from the first topological sorting result.

[0163] Next, we will introduce how the first node is determined.

[0164] In this embodiment, the first node can be determined based on the first topological sorting result using a delayed sorting condition.

[0165] The delayed sorting condition is used to indicate the rules for reordering nodes. The delayed sorting condition can be used to filter out nodes that need to be reordered.

[0166] For example, for each node in a neural network, each node can be compared with a delayed sorting condition. If the node meets the delayed sorting condition, it is designated as the first node. If the node does not meet the delayed sorting condition, the next node is checked.

[0167] In the first possible implementation, the nodes in the neural network have different node types. The memory for some node types is reused, while the memory for others is not. For example, when a node's type is Variable, Constant, or Const, data write operations are required, therefore the memory for that node is not reused.

[0168] Memory reuse can refer to multiple nodes sharing memory, meaning that a node's memory space is released when execution reaches its output node, and used to store data from other nodes. Memory non-reuse means that a node's memory space is used only to store its own data.

[0169] For nodes that do not reuse memory, the node's memory is not shared with other nodes. The execution order of the output nodes of that node is adjusted backward, that is, the execution time of the output nodes of that node is delayed. This will only shorten the life cycle of the output nodes of that node and will not affect the life cycle of other nodes.

[0170] For nodes that reuse memory, the node's memory is shared with other nodes. When the execution order of the node's output nodes is adjusted, the node's lifecycle changes, affecting the lifecycle of other nodes that share memory with that node.

[0171] Therefore, the delayed sort condition can indicate nodes that are not reused in memory. Thus, by using the delayed sort condition, nodes that are not reused in memory are selected from the first topological sort result, and the output node of the node that is not reused in memory is determined as the first node. This quickly locates the first node that needs delayed sorting, improving the efficiency of topological sorting.

[0172] In the second possible implementation, the presence of long-lived nodes in the first topological sorting result increases the theoretical minimum memory footprint of the sorting result. Therefore, to reduce neural network memory usage, a delayed sorting condition can be used to indicate a lifespan threshold. This delayed sorting condition identifies long-lived nodes in the first topological sorting result. Thus, nodes whose lifespans satisfy the lifespan threshold indicated by the delayed sorting condition are identified as the first nodes. Here, a long-lived node can refer to a node whose lifespan satisfies the lifespan threshold indicated by the delayed sorting condition. Alternatively, a node whose lifespan satisfies the lifespan threshold indicated by the delayed sorting condition can refer to a node whose lifespan is greater than the lifespan threshold indicated by the delayed sorting condition, or a node whose lifespan is less than the lifespan threshold indicated by the delayed sorting condition.

[0173] In the first example, the lifecycle threshold indicated by the delayed sorting condition can be the lifecycle of the output node of the node. The method for setting the lifecycle threshold can be found in the description of the lifecycle threshold setting method in S330 above, and will not be repeated here.

[0174] When the lifespan threshold is the lifespan of the output node of a node, the first node is the node whose lifespan is greater than the lifespan threshold indicated by the delayed sorting condition.

[0175] For example, for each node in the first topological sorting result, the lifespan of each node can be compared with the lifespan of its output node. If the lifespan of the node is greater than the lifespan of its output node, then the node is determined to be the first node. If the lifespan of the node is not greater than the lifespan of its output node, then the next node is checked.

[0176] In the second example, when there are multiple output nodes, the node's lifetime can refer to the maximum lifetime of the node's memory. Correspondingly, the lifetime threshold indicated by the lazy sorting condition can refer to the shortest lifetime among the lifetimes of the node's multiple output nodes.

[0177] When the lifecycle threshold is the shortest lifecycle among multiple output nodes, the first node is the node whose lifecycle is greater than the lifecycle threshold indicated by the delayed sorting condition.

[0178] For example, the maximum lifetime of each node is compared with the shortest lifetime of its output node. If the maximum lifetime of a node is greater than the shortest lifetime of its output node, then that node is designated as the first node.

[0179] The maximum lifetime can refer to the maximum length among the distances between a node and each of its output nodes. Similarly, the minimum lifetime of an output node is similar to the maximum lifetime of a node; when there are multiple output nodes, the shortest lifetime is selected from the lifetimes of all the node's output nodes. Alternatively, when an output node corresponds to multiple output nodes, the minimum length among the distances between the output node and each of its corresponding output nodes is taken as the shortest lifetime.

[0180] For example, with Figure 1 Taking the diagram in (c) as an example, if node A's output nodes include nodes B and C, then node A's maximum generation period is [1,3]. For example, let's take... Figure 5 For example, the output nodes of node B include node C and node D. Since the execution order of node C is 5 and the execution order of node D is 3, the maximum lifespan of node B is [2,5].

[0181] In the third example, the lifetime threshold indicated by the delayed sorting condition can be the lifetime of the node's input node.

[0182] When the lifespan threshold is the lifespan of the input node of a node, the first node is the node whose lifespan is less than the lifespan threshold indicated by the delayed sorting condition.

[0183] For example, for each node in the first topological sorting result, the lifespan of each node can be compared with the lifespan of its input node. If the lifespan of the node is less than the lifespan of its input node, then the node is determined to be the first node. If the lifespan of the node is not less than the lifespan of its input node, then the next node is checked.

[0184] In the fourth example, when there are multiple input nodes, the lifetime threshold indicated by the delayed sorting condition can refer to the longest lifetime among the lifetimes of the node's multiple input nodes.

[0185] When the lifespan threshold is the longest lifespan among the lifespans of multiple input nodes of a node, the first node is the node whose lifespan is less than the lifespan threshold indicated by the delayed sorting condition.

[0186] In the third possible implementation, if lifecycle identification is performed on each node in the first topology sorting result, the topology sorting will take a significant amount of time when there are many nodes in the neural network. Furthermore, if the first node is identified only based on node type or lifecycle, there is a risk of missed identifications, thus failing to effectively reduce the memory usage of the neural network. Therefore, delayed sorting conditions can indicate whether a node's memory is reused and a node's lifecycle threshold. In this way, by using node type and node lifecycle, the output nodes of nodes whose memory is not reused and nodes with long lifecycles are determined from the first topology sorting result.

[0187] In one example, for the second node identified as a memory reuse node, the second possible implementation for determining the first node can be referenced to identify whether the second node is a long-lived node. If the second node is a long-lived node, it is identified as the first node. If the memory reuse node is not a long-lived node, the next second node is detected. Here, the second node is the node in the neural network whose type is memory reuse.

[0188] In another example, referring to the second possible implementation for determining the first node described above, we identify whether a node in the first topological sorting result is a long-lived node. If the node is a long-lived node, it is determined as the first node. If the node is not a long-lived node, referring to the first possible implementation for determining the first node described above, if the node is a node that reuses memory, we check the next node. If the node is a node that does not reuse memory, its output node is determined as the first node.

[0189] In this embodiment of the application, after determining the first node, the second execution order of the first node is set according to the output node corresponding to the first node.

[0190] Next, we will take moving the execution order of nodes backward as an example to introduce how to set the second execution order of the first node.

[0191] In the first possible implementation, the lifecycle of the first node can be compared with the lifecycle of its corresponding output node. If the lifecycle of the first node is longer than that of the output node, the first execution order of the output node's predecessor node is used as the second execution order of the first node. If the lifecycle of the first node is not longer than that of the output node, the first execution order of the first node is not adjusted; that is, the first execution order of the first node is used as the second execution order.

[0192] The lifecycle of the output node can be either the shortest or the longest possible lifecycle; however, this embodiment does not impose any limitations on this.

[0193] In one example, if there are other nodes besides the first node in the branch where the first node is located, the first execution order of the nodes in the branch where the first node is located, which are after the first execution order of the first node, can be moved to the back. That is, the first execution order of all nodes in the branch where the first node is located is moved to before the output node of the first node.

[0194] like Figure 6 As shown, Figure 6 This is a flowchart illustrating the process of determining the second execution order of a first node according to an embodiment of this application. The flowchart shown includes steps S331 to S334.

[0195] S331, Create a node sequence based on each node in the first topological sorting result.

[0196] The node sequence comprises multiple subsequences, each containing at least one node. In a DAG of a neural network that contains no two branches, or in a DAG where only one branch contains multiple nodes, each subsequence contains one node. In a DAG of a neural network that contains multiple branches, and at least two branches containing multiple nodes exist, the node sequence contains at least one subsequence containing multiple nodes.

[0197] In one example, an initial node sequence can be created. This initial node sequence is empty. The first execution order of each node in the first topological sort result is traversed, and the nodes are placed into their corresponding subsequences within the initial node sequence according to their first execution order, thus obtaining the node sequence.

[0198] For example, in the case where each subsequence in the node sequence contains one node, the result of the first topological sort is: Figure 1 As shown in Figure (c), Figure 7As shown, the first execution order of each node in the first topological sort result is traversed. Since node B's branch contains only node B, each node corresponds to a subsequence, and the created node sequence is as follows. Figure 7 As shown in Figure (a), each subsequence in the node sequence contains one node.

[0199] For example, in the case where a subsequence in the node sequence contains multiple nodes, the result of the first topological sort is... Figure 5 As shown, Figure 7 As shown, the first execution order of each node in the first topological sort result is traversed. Since the branch containing node C includes nodes C and F, and nodes C and F correspond to the same subsequence, nodes C and F are written into the subsequence of node C in the order of node C→node F, creating a sequence as shown. Figure 7 The node sequence shown in Figure (b) is shown. Subsequence 3 contains nodes C and F.

[0200] S332, when the subsequence containing the first node contains a node, compare the lifetime of the first node with the lifetime of the output node.

[0201] S333: If the lifetime of the first node is longer than the lifetime of the output node, the first execution order of the output node's predecessor node shall be used as the second execution order of the first node.

[0202] In one possible implementation, the first node in the node sequence can be designated as the head node of the subsequence containing the output node, while simultaneously clearing the subsequence containing the first node. In this way, the first execution order of the output node's preceding nodes becomes the second execution order of the first node. Figure 1 Taking the first topological sorting result shown in Figure (c) as an example, when node B is the first node and node F is the output node corresponding to the first node, in Figure 7 In the node sequence shown in Figure (a), node B is added to subsequence 4, where node F is located, and becomes the head node of subsequence 4. Simultaneously, subsequence 2 is cleared, as shown below. Figure 8A As shown in Figure (a).

[0203] S334, when the subsequence containing the first node contains multiple nodes, the first execution order of the preceding nodes of the output node is taken as the second execution order of the first node.

[0204] In a first possible implementation, using the first execution order of the output node's preceding nodes as the second execution order of the first node can include: making the first node the head node of the subsequence containing the output node corresponding to the first node, while simultaneously deleting the first node from the subsequence containing the first node. Figure 7 Taking the node sequence shown in Figure (b) as an example, when the first node is node C, node C is added to subsequence 5, where node G is located, and becomes the head node of subsequence 5. Simultaneously, the first node in subsequence 3 is deleted, at which point subsequence 3 contains only node F. Figure 8A As shown in Figure (b). Then, node F in subsequence 3 is added as the new first node to subsequence 5, where node G is located, and becomes the new tail node of node C in subsequence 5. Simultaneously, node F in subsequence 3 is deleted, as shown in Figure (b). Figure 8A As shown in Figure (c).

[0205] In two possible implementations, using the first execution order of the output node's preceding nodes as the second execution order of the first node can include: moving the subsequence containing the first node to the head of the subsequence containing the output node corresponding to the first node, forming a new subsequence, thereby using the first execution order of the output node's preceding nodes as the second execution order of the first node. In other words, the first execution order of all nodes contained in the branch containing the first node is delayed. For example, as... Figure 8A As shown in Figure (a), when node F is the first node, since the subsequence where node F is located also contains node B, when it is necessary to adjust the execution order of node F, node B → node F can be directly added to the subsequence where the output node corresponding to node F is located, and node B → node F can be used as the head node of the subsequence where the output node corresponding to node F is located.

[0206] Moving the subsequence containing the first node to the head of the subsequence containing the output node corresponding to the first node can be done by: taking each node in the subsequence containing the first node as the head node in the subsequence containing the output node corresponding to the first node according to the last-to-first order of the first execution order, and moving the nodes in the subsequence containing the first node to the subsequence containing the output node corresponding to the first node in turn.

[0207] In one example, if there are other branches in the branch containing the first node in the DAG graph of the neural network, it is necessary to further determine the execution order of the nodes contained in each branch in the branch containing the first node. Therefore, before moving the subsequence containing the first node to the head of the subsequence containing the output node corresponding to the first node, it is necessary to determine whether there are other branches in the branch containing the first node in the DAG graph.

[0208] In a DAG (Directed Acyclic Graph), the statement that "no other branches exist within the branch containing the first node" can mean that all nodes in that branch have a single input and a single output. Conversely, the statement that "other branches exist within the branch containing the first node" can mean that not all nodes in that branch have a single input and a single output.

[0209] For example, in a DAG of a neural network, if all nodes in the branch containing the first node have single inputs and single outputs, the subsequence containing the first node is moved to the head of the subsequence containing the output node of the first node. If, in a DAG of a neural network, the nodes in the branch containing the first node do not all have single inputs and single outputs, the first execution order of the first node is not adjusted; instead, the next node is checked.

[0210] Next, taking the example of moving the execution order of nodes forward, we will introduce how to set the second execution order of the first node.

[0211] In the first possible implementation, similar to shifting the execution order of nodes backward, if the lifetime of the first node is shorter than the lifetime of the input node, the first execution order of the subsequent nodes of the input node is taken as the second execution order of the first node. If the lifetime of the first node is not shorter than the lifetime of the output node, the first execution order of the first node is not adjusted.

[0212] Similar to shifting the execution order of nodes backward, see the above. Figure 6 In the provided embodiment, if there are other nodes besides the first node in the branch where the first node is located, the first execution order of the nodes in the branch where the first node is located that are after the first execution order of the first node can be moved forward, that is, the first execution order of all nodes in the branch where the first node is located can be moved as a whole after the input node of the first node.

[0213] For example, with Figure 7 Taking the node sequence shown in Figure (a) as an example, when the first node is node C, node A is the input node of node C. When moving the execution order of node C forward, nodes C, D, and C can be added to the end of subsequence 1 of node A in sequence, while subsequence 3 is cleared, as shown in Figure (a). Figure 8B As shown in Figure (a), the subsequence 1 containing node A includes node A→node C→node D→node E.

[0214] For example, with Figure 7Taking diagram (b) as an example, when the first node is node D, node B is the input node of node D. Node D is added to subsequence 2 where node B is located, and becomes the tail node of subsequence 2. Simultaneously, node D is deleted from subsequence 4, as shown below. Figure 8B As shown in Figure (b). Then, node E in subsequence 4 is taken as the new first node, node E is added to subsequence 2 where node B is located, and becomes the new tail node in subsequence 4. At the same time, node B in subsequence 4 is deleted, as shown in Figure (b). Figure 8B As shown in Figure (c).

[0215] In this embodiment of the application, after determining the first execution order of the first node, the first topology sorting result can be updated according to the first execution order of the first node as described in S340 above, to obtain the second topology sorting result.

[0216] In one possible implementation, after adjusting the first execution order of the first node through the node sequence, and after setting the first node as the head node in the subsequence containing the output node corresponding to the first node, and clearing the subsequence containing the first node, all nodes in the node sequence can be traversed, and the nodes in the non-empty subsequences of the node sequence can be sorted to obtain the node sequence after the order is rearranged. Based on the index of the nodes in the node sequence after the order is rearranged, the second topological sorting result is obtained.

[0217] For example, with Figure 7 Taking the node sequence shown in Figure (a) as an example, according to Figure 8A As shown in Figure (a), the nodes in the non-empty subsequences of the node sequence are reordered to obtain the following result: Figure 9 The node sequence after rearrangement is shown in Figure (a). The resulting second topological sort is {A,C,D,E,B,F}. For example, using... Figure 7 Taking the node sequence shown in Figure (b) as an example, according to Figure 8A The node sequence shown in Figure (c) is reordered by sorting the nodes in the non-empty subsequences of the node sequence, resulting in the following: Figure 9 The node sequence after rearrangement is shown in Figure (b). The resulting second topological sort is {A,B,D,E,B,F,G}.

[0218] For example, with Figure 7 Taking the node sequence shown in Figure (a) as an example, according to Figure 8B As shown in Figure (a), the nodes in the non-empty subsequences of the node sequence are reordered to obtain the following result: Figure 9 The node sequence after rearrangement is shown in Figure (a). For example, with... Figure 7 Taking the node sequence shown in Figure (b) as an example, according to Figure 8BThe node sequence shown in Figure (c) is reordered by sorting the nodes in the non-empty subsequences of the node sequence, resulting in the following: Figure 9 The node sequence after the order is rearranged is shown in Figure (b).

[0219] For example, consider moving the execution order of nodes backwards, such as... Figure 10 As shown, Figure 10 This is a flowchart illustrating the process of determining the second topological sorting result according to an embodiment of this application. Each node in the DAG of the neural network is sorted using a topological sorting algorithm to obtain a first topological sorting result (S11). Each node in the first topological sorting result is traversed, and each node is placed into its subsequence to create a node sequence (S12). For the nth node, it is determined whether node n is the first node (S13). If node n is not the first node, the (n+1)th node is obtained (S18). If node n is the first node, it is determined whether the subsequence containing node n contains a node (S14). If the subsequence containing node n contains a node, it is determined whether node n is a long-lived node (S15). If node n is not a long-lived node, S18 is executed to obtain the (n+1)th node. If node n is a long-lived node, the output node m of node n is obtained from the node sequence, node n is added to the head of the subsequence of node m, and the subsequence containing node n is cleared (S16). If the subsequence containing node n contains multiple nodes, determine whether node n is a single-input, single-output node (S17). If node n is a single-input, single-output node, proceed to S16. If node n is not a single-input, single-output node, proceed to S18 and select the (n+1)th node. Repeat the above operation until n is greater than N, i.e., when all nodes in the node sequence have been detected, sort the nodes in the non-empty subsequences of the node sequence to obtain the second topological sorting result (S19).

[0220] Where n is an integer greater than or equal to 1 and less than or equal to N, and m is an integer greater than 1 and less than or equal to N, where m is greater than n. N indicates the number of nodes in the node sequence. In this embodiment, after obtaining the second topological sorting result, memory allocation for nodes in the neural network can be performed based on the second topological sorting result. When the neural network processes the data to be processed, the processing operation is performed according to the second execution order of the nodes in the second topological sorting result to obtain the processing result of the data to be processed.

[0221] The data to be processed includes at least one of text data, image data, or video data, and this application embodiment does not limit this.

[0222] The processing operations include at least one of the following: recognition operations, content generation operations, or feature extraction operations.

[0223] The identification operation can refer to identifying the content of the data or identifying the type of data.

[0224] For example, taking image data as an example, the content of the data to be processed can be, for example, object detection or object recognition. The type of data to be recognized can be, for example, image classification.

[0225] For example, taking text data as an example, content generation operations can refer to text style conversion, text summarization generation, text-to-image conversion, etc.

[0226] Content generation refers to the automatic generation of content such as text, images, and videos using artificial intelligence algorithms. For example, content generation can be performed using generative models. A generative model can be a model that generates content by learning the distribution of input data.

[0227] Feature extraction refers to obtaining data features from the data to be processed, such as extracting semantic features from text, image features, etc.

[0228] In one possible implementation, allocating memory for nodes in the neural network based on the second topological sorting result may include: determining the lifecycle of nodes in the DAG of the neural network according to the second topological sorting result; and allocating memory locations to nodes in the DAG of the neural network according to the lifecycle of nodes in the DAG.

[0229] In this context, memory location can refer to the storage address of the memory in the computing device that runs the neural network.

[0230] For example, nodes with non-overlapping lifecycles in a neural network's Directed Acyclic Graph (DAG) can be identified based on their lifecycles. Memory locations can then be allocated to these nodes using memory reuse techniques. In this way, by allocating memory locations to nodes with non-overlapping lifecycles, the memory footprint of the neural network can be reduced.

[0231] For example, taking the second topological sort result as Figure 1 Taking the diagram in (d) as an example, the node execution order is {A, C, D, E, B, F}. Nodes B, C, and D have non-overlapping lifecycles. When all nodes occupy a single memory block, the size of that memory block is 1MB. Figure 11As shown, memory block 1 is allocated to node B. Since the lifetime of node A is [1,5] and the lifetime of node B is [5,6], node B and node A share memory, meaning that node A and node B share memory block 1 at different times. Memory block 2 is allocated to node C. Since the output of node C is input to node D, and the output of node D is input to node E, the memory of nodes C, D, and E is shared, meaning that nodes C, D, and E share memory block 2 at different times. In this way, the memory usage of the neural network is reduced to two memory blocks.

[0232] The above primarily uses AI platform 101 as an example to introduce the technical solutions provided in the embodiments of this application. It is understood that the functions performed by AI platform 101, as an example, include the corresponding hardware structures and / or software modules for executing each functional module. Those skilled in the art should readily recognize that, in conjunction with the unit and algorithm operations of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0233] This application embodiment can divide the AI ​​platform 101 into functional modules according to the above method example. Each function can be divided into separate functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It is understood that the naming and grouping of devices and modules in this application embodiment are illustrative and represent only one logical functional grouping; in actual implementation, other grouping methods may be used.

[0234] For example, the deep learning framework layer 1012 in the AI ​​platform 101 can be named the topological sorting device 12. Figure 12 As shown, Figure 12 This is a schematic diagram of the topology sorting device provided in an embodiment of this application. The topology sorting device 12 shown includes a sorting module 121, a node determination module 122, an order correction module 123, and a topology determination module 124.

[0235] The sorting module 121 is used to obtain a first topological sorting result. The first topological sorting result is used to indicate the first execution order of nodes in the neural network.

[0236] The node determination module 122 is used to determine the first node based on the first topological sorting result. The first node is a node that has been reordered in the neural network.

[0237] The sequence correction module 123 is used to update the first execution order of the first node to obtain the second execution order of the first node. The lifespan of the first node in the second execution order is shorter than the lifespan of the first node in the first execution order.

[0238] The topology determination module 124 is used to update the first topology sorting result according to the second execution order of the first node to obtain the second topology sorting result.

[0239] In one possible implementation, the topology sorting device 12 may further include a memory allocation module and a model execution module. Figure 12 (not shown in the image), etc.

[0240] The memory allocation module is used to allocate memory locations of computer-accessible memory to nodes of the neural network based on the second topology sorting result.

[0241] The model execution module is used to process the input data using a computer and output the processing results. The data to be processed includes at least one of text data, image data, or video data. The processing results include at least one of recognition results or generated content.

[0242] The sorting module 121, node determination module 122, order correction module 123, topology determination module 124, memory allocation module, and model execution module can all be implemented in software or hardware. The implementation of the sorting module 121 will be described below using it as an example. Similarly, the implementation methods of the node determination module 122, order correction module 123, topology determination module 124, memory allocation module, and model execution module can refer to the implementation method of the sorting module 121.

[0243] As an example of a software functional unit, the sorting module 121 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the sorting module 121 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0244] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0245] As an example of a hardware functional unit, the sorting module 121 may include at least one computing device, such as a server. Alternatively, the sorting module 121 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0246] The sorting module 121 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the sorting module 121 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the sorting module 121 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0247] It should be noted that, in other embodiments, the sorting module 121, the node determination module 122, the order correction module 123, and the topology determination module 124 can all execute any step in the topology sorting method. The steps implemented by the sorting module 121, node determination module 122, order correction module 123, topology determination module 124, memory allocation module, and model execution module can be specified as needed. The sorting module 121, node determination module 122, order correction module 123, topology determination module 124, memory allocation module, and model execution module respectively implement different steps in the topology sorting method to achieve all the functions of the topology sorting device 12.

[0248] This application also provides a computing device for implementing the above-described topological sorting method.

[0249] In one example, the computing device includes Figure 12 The topology sorting device 12 in the middle. The topology sorting device 12 includes a sorting module 121, a node determination module 122, an order correction module 123 and a topology determination module 124.

[0250] In another example, the computing device may include, for example: Figure 2 The AI ​​platform 101 shown.

[0251] In another example, such as Figure 13 As shown, the computing device 13 may include a bus 132, a processor 134, a memory 136, and a communication interface 138. The processor 134, the memory 136, and the communication interface 138 communicate with each other via the bus 132. The computing device 13 may be a server or a terminal device. It should be understood that this application does not limit the number of processors 134 and memory 136 in the computing device 13.

[0252] Bus 132 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 13 The bus 132 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 132 may include a path for transmitting information between various components of the computing device 13 (e.g., memory 136, processor 134, communication interface 138).

[0253] The processor 134 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0254] In this application, processor 134 performs the above-mentioned... Figure 3 The method is illustrated below. For example, obtain the first topological sorting result, determine the first node based on the first topological sorting result, update the first execution order of the first node to obtain the second execution order of the first node, update the first topological sorting result according to the second execution order of the first node to obtain the second topological sorting result.

[0255] Memory 136 may include volatile memory, such as random access memory (RAM). Processor 134 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0256] The memory 136 stores executable program code, and the processor 134 executes the executable program code to implement the functions of the aforementioned sorting module 121, node determination module 122, order correction module 123, and topology determination module 124, thereby realizing the topology sorting method. That is, the memory 136 stores instructions for executing the topology sorting method.

[0257] In this embodiment of the application, the memory 136 contains repository files, etc.

[0258] The communication interface 138 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 13 and other devices or communication networks.

[0259] The topology sorting method disclosed in the above embodiments can be applied to, or implemented by, processor 134. Processor 134 can be an integrated circuit chip with signal processor capabilities.

[0260] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 134 or by instructions in the form of software. The processor 134 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete vacuum tubes or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of the hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 136, and the processor 134 reads the information in memory 136 and completes the steps of the above method in combination with its hardware.

[0261] In one possible implementation, the processor 134 can also be used to execute a topological sorting method. For specific implementation, please refer to the embodiments provided by the topological sorting method described above. The embodiments of this application will not be repeated here.

[0262] This application also provides a chip system. The chip system includes a processor and input / output ports. The processor is used to implement the processing functions involved in the above-described method embodiments. The input / output ports are used to implement the transmit / receive functions involved in the above-described method embodiments.

[0263] In one possible implementation, the chip system further includes a memory. This memory stores program instructions and data for implementing the functions described in the above method embodiments.

[0264] This chip system can consist of chips or include chips and other discrete components.

[0265] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a terminal of any of the foregoing embodiments, such as an internal storage unit including a data transmission end and / or a data receiving end, like a hard disk or memory of the terminal. The computer-readable storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal. Further, the computer-readable storage medium can include both the internal storage unit and the external storage device of the terminal. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0266] This application provides a computer program product containing instructions, including a computer program or instructions that, when run on a computer, cause the computer to perform the topology sorting method described in the above method embodiments.

[0267] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0268] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0269] It should be noted that the terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0270] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0271] It should be understood that in the embodiments of this application, "B corresponding to A" means that B is associated with A. For example, B can be determined based on A. It should also be understood that determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information. Furthermore, the term "connection" in the embodiments of this application refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices, and the embodiments of this application do not impose any limitations on this.

[0272] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A topological sorting method, characterized in that, The method includes: Obtain the first topological sorting result; the first topological sorting result is used to indicate the first execution order of the nodes in the neural network; The first node is determined based on the first topological sorting result; the first node is a node that has been reordered among the nodes of the neural network. The first execution order of the first node is updated to obtain the second execution order of the first node; the lifecycle of the first node under the second execution order is shorter than the lifecycle of the first node under the first execution order. The first topology sorting result is updated according to the second execution order of the first node to obtain the second topology sorting result.

2. The method according to claim 1, characterized in that, The first node is the node in the first topological sorting result that satisfies the delayed sorting condition; the delayed sorting condition is used to indicate the rules for nodes that satisfy delayed sorting.

3. The method according to claim 2, characterized in that, The delayed sorting condition is used to indicate that the memory of a node is not reused; The step of determining the first node based on the first topological sorting result includes: Select nodes whose memory is not reused from the first topology sorting results, and determine the output node of the node whose memory is not reused as the first node.

4. The method according to claim 2 or 3, characterized in that, The delayed sorting condition is used to indicate the lifecycle threshold corresponding to the node; the lifecycle is used to indicate the length of time between memory allocation and memory release. The step of determining the first node based on the first topological sorting result includes: Select the first node from the first topology sorting results whose life cycle satisfies the life cycle threshold corresponding to the node.

5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the first node based on the first topological sorting result includes: From the second node, select the first node whose lifecycle satisfies the lifecycle threshold corresponding to the node; the second node is the node whose memory is reused in the first topology sorting result.

6. The method according to claim 4 or 5, characterized in that, The lifecycle threshold corresponding to the node is the lifecycle of the node's output nodes, or the lifecycle threshold corresponding to the node is the shortest lifecycle among the lifecycles of the node's multiple output nodes.

7. The method according to any one of claims 1 to 6, characterized in that, The step of updating the first execution order of the first node to obtain the second execution order of the first node includes: Based on the output node of the first node, update the first execution order of the first node to obtain the second execution order of the first node; the interval between the first node and the output node in the second execution order is less than the interval between the first node and the output node in the first execution order.

8. The method according to claim 7, characterized in that, The step of updating the first execution order of the first node based on the output node of the first node to obtain the second execution order of the first node includes: Compare the lifecycle of the first node with the lifecycle of the output node; If the lifecycle of the first node is longer than the lifecycle of the output node, the first execution order of the preceding node of the output node is taken as the second execution order of the first node; the preceding node is the node whose first execution order is adjacent to the first execution order of the output node.

9. The method according to claim 8, characterized in that, Before comparing the lifecycle of the first node with the lifecycle of the output node, the method further includes: Based on each node in the first topological sorting result, a node sequence is created; the node sequence includes multiple subsequences, each node corresponds to one subsequence, and each subsequence contains at least one node; When the subsequence containing the first node contains a node, the lifetime of the first node is compared with the lifetime of the output node.

10. The method according to claim 9, characterized in that, After creating a node sequence based on each node in the first topological sorting result, the method further includes: When the subsequence containing the first node contains multiple nodes, the first execution order of the preceding node of the output node is taken as the second execution order of the first node; the preceding node is the node whose first execution order is adjacent to the first execution order of the output node.

11. The method according to any one of claims 1 to 6, characterized in that, The step of updating the first execution order of the first node to obtain the second execution order of the first node includes: Based on the input node of the first node, the first execution order of the first node is updated to obtain the second execution order of the first node; the interval between the first node and the input node in the second execution order is less than the interval between the first node and the input node in the first execution order.

12. The method according to claim 7, characterized in that, The step of updating the first execution order of the first node based on the input node of the first node to obtain the second execution order of the first node includes: Compare the lifecycle of the first node with the lifecycle of the input node; If the lifecycle of the first node is shorter than the lifecycle of the input node, the first execution order of the subsequent nodes of the input node is taken as the second execution order of the first node; the subsequent node is the node whose first execution order is adjacent to the first execution order of the input node.

13. The method according to any one of claims 1 to 12, characterized in that, Obtaining the first topological sort result includes: Configure the execution order of each node in the directed acyclic graph (DAG) of the neural network to obtain the first topological sorting result.

14. The method according to any one of claims 1 to 13, characterized in that, After updating the first topology sorting result according to the second execution order of the first node to obtain the second topology sorting result, the method further includes: Based on the second topology sorting result, memory locations of computer-accessible memory are allocated to the nodes of the neural network; The data to be processed is input into the computer; the data to be processed includes at least one of text data, image data, or video data. The computer is used to process the data to be processed and output the processing result of the data to be processed; the processing result includes at least one of recognition result or generated content.

15. The method according to claim 14, characterized in that, The step of allocating computer-accessible memory locations to nodes of the neural network based on the second topological sorting result includes: Based on the second topological sorting result, the lifecycle of the nodes in the neural network is determined; The memory location is allocated to the nodes of the neural network according to the lifecycle of the nodes in the DAG of the neural network.

16. A topological sorting device, characterized in that, The device includes: A sorting module is used to obtain a first topological sorting result; the first topological sorting result is used to indicate the first execution order of nodes in the neural network. A node determination module is used to determine a first node based on the first topological sorting result; the first node is a node that has been reordered among the nodes of the neural network. The sequence correction module is used to update the first execution order of the first node to obtain a second execution order of the first node; the lifespan of the first node in the second execution order is shorter than the lifespan of the first node in the first execution order. The topology determination module is used to update the first topology sorting result according to the second execution order of the first node to obtain the second topology sorting result.

17. An artificial intelligence (AI) platform, characterized in that, The AI ​​platform is used to provide users with services for building and running neural networks; The AI ​​platform is used to allocate memory locations of computer-accessible storage to nodes of the neural network based on the topological sorting results of the nodes in the neural network; and to input the data to be processed into the computer. The data to be processed includes at least one of text data, image data, or video data; the computer is used to process the data to be processed; the processing operation is used to identify or generate at least one of these; the computer is used to output the processing result of the data to be processed. The topological sorting result is obtained by the topological sorting method described in any one of claims 1 to 12.

18. A computing device, characterized in that, The computing device includes: a processor coupled to a memory; The memory is used to store computer programs; The memory is configured to execute the computer program stored in the memory, so that the computing device performs the method as described in any one of claims 1 to 15.

19. A computer program product, characterized in that, The computer program product includes: a computer program or instructions that, when run on a computing device, cause the computing device to perform the method as described in any one of claims 1 to 15.