System and method for avoiding memory demand surge in multi-stream parallel data processing

By reconstructing the tensor release node in multi-stream parallel data processing, ensuring that the tensor is released only after all dependent nodes are used, the problem of skyrocketing memory demand is solved and system efficiency is improved.

CN114860459BActive Publication Date: 2025-05-20BEIJING SILICONFLOW TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210642068.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-05-20
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

In multi-stream parallel data processing, parameter exchange and data dependence between task flows lead to a surge in memory demand, which in turn affects system efficiency.

Method used

When inserting the computing logic node, the task flow calculation engine confirms whether it is a tensor-generated logical node, and uses the tensor information acquisition component and the tensor release node to reconstruct the component, reconstruct the tensor-free node, so that it is connected to the tensor generation and use nodes, ensuring that the tensor is released only after all dependent nodes are used, thereby controlling the rhythm of memory application.

Benefits of technology

It effectively avoids the surge in memory demand, prevents the surge in memory applications caused by the rapid processing of tasks with fast computing speed, ensures the system's reasonable use of memory, and improves the system efficiency of multi-stream parallel data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860459B_ABST
    Figure CN114860459B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a system and method for avoiding a surge in memory demand in multi-stream parallel data processing. The system includes: a task flow computing engine, which, as the task is executed, confirms whether the computing logic node to be inserted is a tensor generation logic node when inserting the computing logic node task contained in the task into the corresponding task flow to which it belongs; a tensor information acquisition component, which, when the task flow computing engine confirms that the computing logic node to be inserted is not a tensor generation logic node, obtains information of all computing logic nodes that need to use the generated tensor; and a tensor release node reconstruction component, which reconstructs the tensor release node based on the information obtained by the tensor information acquisition component, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow to which it belongs, so that the subsequent computing logic nodes in the task flow where the reconstructed tensor release node is located apply for memory execution after the tensor release node completes the tensor release.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data processing technology. More specifically, the present disclosure relates to a system and method for avoiding a sharp increase in memory requirements in stream parallel data processing. Background Art

[0002] Nowadays, with the popularization of deep learning, in order to improve the speed of data processing, it is usually necessary to divide a task into multiple shard tasks to form multiple task streams, so that the task streams perform parallel processing, thereby saving the time of data processing or accelerating the efficiency of task processing. This data processing method is multi-stream parallel processing.

[0003] However, in the multi-stream parallel data processing method, there are often parameter exchanges between different streams and data dependencies between different streams. Therefore, if parameter synchronization is not achieved in parallel, errors will occur when the results generated by each task stream are merged. Since there may be tensor dependencies among different computing logic nodes of each task stream, when there are differences in the data processing speeds of different task streams, the fast task stream will continuously apply for the memory required for computing, resulting in a sharp increase in the memory requirements of the entire system. And due to the dependence of the slow task stream on the parameters in the fast task stream, although the fast task stream applies for a large amount of memory, it is always in a state where the memory cannot be released, resulting in a lack of available memory in the data processing system. When other task streams need to apply for memory, they will be blocked, resulting in a reduction in the efficiency of the entire parallel processing system.

[0004] Therefore, there is a need for a data processing system and method that can eliminate the sharp increase in memory requirements when implementing multi-stream parallelism. Summary of the Invention

[0005] An object of the present invention is to solve at least the above problems. Specifically, the present disclosure provides a system for avoiding a sharp increase in memory requirements in multi-stream parallel data processing, including: a task flow computing engine, which, as a task is executed, when inserting a computational logic node task included in the task into its corresponding task flow, confirms whether the computational logic node to be inserted is a tensor generation logic node. If not, directly insert the computational logic node into its corresponding task flow; a tensor information acquisition component, which, when the task flow computing engine confirms that the computational logic node to be inserted is a tensor generation logic node, acquires information on all computational logic nodes that need to use the generated tensor; and a tensor release node reconstruction component, which, based on the information obtained by the tensor information acquisition component, reconstructs the tensor release node such that the output ends of all computational logic nodes on all task flows that need to use the tensor are connected to the input end of the tensor release node, and the output end of the tensor release node is connected between the computational logic node that uses the tensor and its downstream computational logic nodes in the task flow to which the computational logic node that generates the tensor belongs, so that the task flow computing engine can insert the reconstructed tensor release node into its corresponding task flow, thereby enabling the subsequent computational logic nodes in the task flow where the reconstructed tensor release node is located to apply for memory to be executed after the tensor release node completes tensor release.

[0006] According to the system for avoiding a sharp increase in memory requirements in multi-stream parallel data processing of the present disclosure, wherein, when the tensor is used by other downstream computational logic nodes of the computational logic node that generates the tensor, the task flow computing engine inserts the reconstructed tensor release node immediately before the last computational logic node that uses the tensor in its corresponding task flow.

[0007] According to the system for avoiding a sharp increase in memory requirements in multi-stream parallel data processing of the present disclosure, wherein, after the task flow computing engine emits the reconstructed tensor release node and then receives a release node for the same tensor again, it discards the tensor release node.

[0008] According to another aspect of the present disclosure, a method for avoiding a sharp increase in memory requirements in multi-stream parallel data processing includes: as a task is executed, when a task flow computing engine inserts a computing logic node task included in the task into its corresponding task flow, it confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, it directly inserts the computing logic node into its corresponding task flow; when the task flow computing engine confirms that the computing logic node to be inserted is a tensor generation logic node, it obtains information of all computing logic nodes that need to use the generated tensor through a tensor information acquisition component; and based on the information obtained by the tensor information acquisition component, it reconstructs a tensor release node through a tensor release node reconstruction component, so that the output ends of all computing logic nodes that need to use the tensor on all task flows are connected to the input end of the tensor release node, and the output end of the tensor release node is connected between the computing logic node that uses the tensor and its downstream computing logic nodes in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into its corresponding task flow, so that the subsequent computing logic nodes in the task flow where the reconstructed tensor release node is located apply for memory after the tensor release node completes tensor release.

[0009] The method for avoiding a sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure further includes: when the tensor is used by other downstream computing logic nodes of the computing logic node that generates the tensor, the task flow computing engine inserts the reconstructed tensor release node immediately before the last computing logic node that uses the tensor in its corresponding task flow.

[0010] The method for avoiding a sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure, wherein, after the task flow computing engine emits the reconstructed tensor release node and then receives a release node for the same tensor again, it discards the tensor release node.

[0011] According to another aspect of the present disclosure, there is also provided a system for avoiding an explosive increase in memory requirements in multi-stream parallel data processing, including: a task flow computing engine, which, as the task is executed, when inserting the computing logic node task included in the task into its corresponding task flow, confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, directly insert the computing logic node into its corresponding task flow; a tensor information acquisition component, which, when the task flow computing engine confirms that the computing logic node to be inserted is not a tensor generation logic node, acquires information on all computing logic nodes that need to use the generated tensor; and a tensor release node reconstruction component, which, based on the information obtained by the tensor information acquisition component, reconstructs the tensor release node, such that the output ends of all computing logic nodes on all task flows that need to use the tensor are connected to the input end of the tensor release node, and the output end of the tensor release node is connected to the input end of the downstream computing logic node of the computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow containing the computing logic node that uses the tensor, thereby enabling the execution of memory application by the computing logic node after the last computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs to be completed after the tensor release node completes tensor release.

[0012] According to another aspect of the present disclosure, there is also provided a method for avoiding an explosive increase in memory requirements in multi-stream parallel data processing, including: as the task is executed, when inserting the computing logic node task included in the task into its corresponding task flow through the task flow computing engine, confirm whether the computing logic node to be inserted is a tensor generation logic node. If not, directly insert the computing logic node into its corresponding task flow; when the task flow computing engine confirms that the computing logic node to be inserted is a tensor generation logic node, acquire information on all computing logic nodes that need to use the generated tensor through the tensor information acquisition component; and based on the information obtained by the tensor information acquisition component, reconstruct the tensor release node through the tensor release node reconstruction component, such that the output ends of all computing logic nodes on all task flows that need to use the tensor are connected to the input end of the tensor release node, and the output end of the tensor release node is connected to the input end of the downstream computing logic node of the computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow containing the computing logic node that uses the tensor, thereby enabling the execution of memory application by the computing logic node after the last computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs to be completed after the tensor release node completes tensor release.

[0013] By adopting the system and method for avoiding the sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure, based on the tensor to be processed, by deploying the release node of the tensor in the same task stream as the generation calculation logic node of the tensor, and establishing a direct connection relationship between the release node of the tensor and the calculation logic nodes using the released tensor in other task streams and a direct connection relationship with the downstream calculation logic nodes of the generation calculation logic node, so that the release node of the tensor is reconstructed to form a new tensor release node, enabling the new tensor release node to ensure the timing of tensor release on the one hand, and on the other hand, restricting the rhythm of memory application of the downstream calculation logic nodes in the task stream where it is located, thereby preventing the sharp increase in memory demand caused by the fast processing of the task stream with high computing speed, and eliminating the unlimited requirement of the data processing system for memory.

[0014] Other advantages, objectives and features of the present invention will be partially reflected by the following description, and partially will also be understood by those skilled in the art through the research and practice of the present invention. Brief Description of the Drawings

[0015] Figure 1 The figure shows a schematic diagram of the sharp increase in memory requirements in the existing stream parallel data processing structure.

[0016] Figure 2 The figure shows a schematic diagram of the first embodiment of the system for avoiding the sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure.

[0017] Figure 3 The figure shows a schematic diagram of the second embodiment of the system for avoiding the sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure. Detailed Embodiments

[0018] The following further elaborates on the present invention in conjunction with the embodiments and the drawings, so that those skilled in the art can implement it with reference to the description in the specification.

[0019] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0020] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "that" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0021] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, one of two possible objects may be referred to as the first logical node or the second logical node hereinafter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0022] To enable those skilled in the art to better understand this disclosure, the following further describes this disclosure in detail in conjunction with the accompanying drawings and specific embodiments.

[0023] Figure 1 Shown is a schematic diagram of the soaring memory demand in the existing convection parallel data processing structure.

[0024] As Figure 1 shown, in the existing convection parallel data processing structure, the task flow computing engine launches the computing tasks of the computing logic nodes in each task flow to different task execution flows or execution threads, so that each task flow is executed sequentially. During the execution of the task flow, the computing logic nodes request memory allocation based on their own computing needs, and as the computing tasks are completed, the memory is released and can be applied for by other computing logic nodes. As long as the data stored in the memory applied for by the computing logic node has not been used up or the computing logic node has not completed its own task, the applied memory will not be released. As Figure 1As shown, as the tasks in task flows S1 and S2 are sent to the coprocessors responsible for each task flow, the computing unit makes a memory application. For example, there are any number of computing nodes in task flow S1, such as nodes S1-n1, S1-n, S1-n4, etc. Here, S1-n1 is called the first node. Similarly, there are any number of computing nodes in task flow S2, such as nodes S2-n1, S2-n, S2-n4, etc. Here, S2-n2 is called the second node. There is also a task flow in the CPU, and only a logical node C-rn3 is shown here. This is a tensor release node and can also be other computing logic nodes. Under normal circumstances, the task flow computing engine will send the tensor release node to the specified task flow according to the order of the computational graph of the task, and the release of memory is not subject to any constraints. For example, as Figure 1 shown, when the tensor release node is the logical node C-rn3, if it is launched into the task flow on the CPU, after it is used at T1, it is directly released. To some extent, the nodes after the computing logic node S1-n1 that generates T1 in task flow S1 do not need to know whether T1 is released and can directly enter their execution stage, such as S1-n2, S1-n3, S1-n4, etc., and make continuous memory applications. If various tensors in S1 are generated quickly, and the computing logic nodes in task flow S2, which consume the tensors, are slow, it will cause a large number of computing nodes in S1 to apply for a large amount of memory, eventually leading to a sharp increase in memory demand, resulting in a lack of memory and even affecting the operation of the entire system.

[0025] As Figure 1 shown, taking the tensor T1 as an example, in one case, the tensor release node is the logical node C-rn3 deployed in the task flow S on the CPU. In another case, the tensor release node is directly launched into the execution task flow of S1, such as S1-rn3(T1) represented by a dotted box. In still another case, it is launched after a computing logic node in task flow S2 that finally consumes T1, such as S2-rn3(T1) represented by a dotted box, so as to avoid the memory storing T1 being released in advance, resulting in calculation errors. However, none of these can avoid the situation where, when the processing speed of S1 is fast, the amount of memory it applies for surges.

[0026] Therefore, the present disclosure provides a system 100 for avoiding a sharp increase in memory demand in multi-stream parallel data processing. Figure 2 Shown is a schematic diagram of a first embodiment of a system for avoiding a sharp increase in memory demand in multi-stream parallel data processing according to the present disclosure. As Figure 2As shown. Above it is a computational graph representing the tasks of the program to be executed. The CPU obtains all the computational logic nodes of the computational graph, thereby obtaining the relationships between the computational logic nodes. In the task flow computing engine 130 in the system 100 for avoiding a sharp increase in memory requirements in multi-stream parallel data processing according to the present disclosure, as the task is executed, when inserting the computational logic node task included in the task into its corresponding task flow, it is confirmed whether the computational logic node to be inserted is a tensor generation logic node. If not, the computational logic node is directly inserted into the task flow to be executed to which it belongs.

[0027] When the task flow computing engine 130 confirms that the computational logic node to be inserted is a tensor generation logic node, this means that the release timing of the tensor needs to be rearranged. For this purpose, it is necessary to understand the production and consumption information of the tensor to be released. For this purpose, the tensor information acquisition component 110 acquires the information of all computational logic nodes that need to use the generated tensor. As shown in the figure, in the task flow S1, S1-n2 and in the task flow S2, S2-n2 will both use the tensor T1 generated by the computational logic node S1-n1. For this reason, on the one hand, it is necessary to ensure the correctness of the calculation. After both S1-n2 and S2-n2 have used the tensor T1, the tensor T1 can be released. Therefore, it is necessary to determine that the memory release logic node for releasing the tensor T1 is restricted by S1-n2 and S2-n2. Therefore, the information related to the tensor acquired by the tensor information acquisition component 110 includes: the logic node that generates the tensor, the direct downstream logic node of the logic node that generates the tensor, and the logic node that uses the tensor, and the computational logic nodes that use the tensor in other task flows.

[0028] Subsequently, regardless of whether there is a memory release logic node for the tensor in the initial task-corresponding computational graph, the tensor release node reconstruction component 120 reconstructs the tensor release node based on the information obtained by the tensor information acquisition component 110. If there is a tensor release node for the tensor, the tensor release node reconstruction component 120 reconstructs the tensor release node by rewriting based on the information obtained by the tensor information acquisition component 110. If there is no tensor release node for the tensor, the tensor release node reconstruction component 120 directly constructs a tensor release node based on the information obtained by the tensor information acquisition component 110. For example, in the case where there is a tensor release node for the tensor T1, under normal circumstances, it will be directly sent to the task flow S1 or task flow S2 as shown in Figure 1 shown or retained in the task flow S processed by the CPU itself. But in the case as shown in Figure 2 shown, the tensor release node reconstruction component 120 will reconstruct, for the tensor T1, based on the relationship between the tensor and the computational logic nodes S1-n1, S1-n2, and S2-n2, as shown in Figure 2The tensor release node S1-rn3(T1) shown is located after the downstream computing logic node S1-n2 of the computing logic node S1-n1 that generates the tensor T1, and enables an execution control connection relationship between the tensor release node S1-rn3(T1) and the computing logic node S2-n2 that uses the tensor T1 in the task flow S2, that is, the execution body corresponding to the computing logic node S2-n2 will send a status message to its downstream after completing the use of the tensor T1, expressing that it has finished using the tensor T1. In this way, the tensor release node S1-rn3(T1) needs to wait for all the computing logic nodes that use the tensor T1 in the task flow S1 generated by the tensor T1 to be released before it can perform the release operation, thereby ensuring that all computing logic nodes will not perform operations on the wrong tensor. In addition, in order to limit the memory application of the subsequent computing logic nodes in the task flow S1, the tensor release node S1-rn3(T1) is set before the downstream computing logic node (for example, S1-n4) of the computing logic node that last uses tensor T1 in S1, so that the input and output ends of the tensor release node S1-rn3(T1) are connected to the input end of the downstream logic node, so that the execution body corresponding to the downstream computing logic node needs to receive the message that the execution body corresponding to the tensor release node S1-rn3(T1) has completed the release of tensor T1 before it can start to execute operations and apply for memory. The tensor release node S1-rn3(T1) is used to limit the rhythm of subsequent computing logic nodes to apply for memory to control the surge in memory demand of task flow S1 when the processing speed is relatively fast. Therefore, the tensor release node reconstruction component 120 connects the output ends of the computing logic nodes that need to use the tensor on all task flows to the input end of the tensor release node, and connects the output end of the tensor release node to the computing logic node that uses the tensor and its downstream computing logic node in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine 110 inserts the reconstructed tensor release node into the task flow to which it belongs, so that the subsequent computing logic nodes in the task flow where the reconstructed tensor release node is located apply for memory execution after the tensor release node completes the tensor release.

[0029] Figure 3 Shown is a schematic diagram of a second embodiment of a system for avoiding skyrocketing memory requirements in multi-stream parallel data processing according to the present disclosure. Figure 3 The system shown is similar to Figure 2 The system structure is the same as shown in Figure 3 In the case shown in , the tensor release node reconstruction component 120 reconstructs the tensor T1 based on the relationship between the tensor and the computing logic nodes S1-n1, S1-n2 and S2-n2 as shown in Figure 3The shown tensor release node S2-rn3(T1) is such that the output ends of the computational logic nodes that finally use the tensor T1 in all task flows are connected to the input end of the tensor release node S2-rn3(T1), and the output end of the tensor release node S2-rn3(T1) is connected to the input end of the direct downstream computational logic node (such as S1-n4) of the computational logic node that finally uses the tensor T1 in the task flow S1, thereby establishing a control connection edge between the tensor release node S2-rn3(T1) and the computational logic node S1-n4. In this way, the execution body corresponding to the computational logic node S2-n2 will send a status message to its downstream after completing the use of the tensor T1, indicating that it has finished using the tensor T1. In this way, the tensor release node S2-rn3(T1) needs to wait for all the computational logic nodes that use the tensor T1 in all task flows to complete the use before it can perform the release operation, thereby ensuring that all computational logic nodes do not operate on the wrong tensor. In addition, in order to limit the memory application of the subsequent computational logic nodes in the task flow S1, before connecting the outlet end of the tensor release node S2-rn3(T1) to the downstream computational logic node (such as S1-n4) of the computational logic node that finally uses the tensor T1 in S1, the input-output end of the tensor release node S2-rn3(T1) is connected to the input end of this downstream logic node, so that the execution body corresponding to this downstream computational logic node needs to receive the message that the execution body corresponding to the tensor release node S2-rn3(T1) has completed the release of the tensor T1 before it can start to execute the operation and apply for memory. By using the tensor release node S2-rn3(T1) to limit the rhythm of memory application of the subsequent computational logic nodes, the soaring demand for memory in the task flow S1 in the case of relatively fast processing speed can be controlled. Therefore, the tensor release node reconstruction component 120 makes the output ends of all the computational logic nodes that need to use the tensor on all task flows be connected to the input end of the tensor release node, and makes the output end of the tensor release node be connected to the input end of the downstream computational logic node of the computational logic node that uses the tensor in the task flow to which the computational logic node that generates the tensor belongs, so that the task flow computing engine 110 inserts the reconstructed tensor release node into the task flow containing the computational logic node that uses the tensor, so that the execution of memory application by the computational logic nodes after the computational logic node that finally uses the tensor in the task flow to which the computational logic node that generates the tensor belongs is after the tensor release node completes the tensor release.

[0030] By reconstructing the tensor release node as described above, especially reconstructing the association relationship between the tensor release node and the node that consumes the tensor, especially connecting the input end of the downstream computational logic node of the computational logic node that finally uses the tensor in the task flow that generates the tensor to the output end of the reconstructed tensor release node, the situation of soaring memory application in the task flow where the tensor generation node is located can be controlled.

[0031] Alternatively, the reconstructed tensor release node may also be deployed in the task flow S run by the CPU, as long as the output ends of all computational logic nodes that need to use the tensor on all task flows are connected to the input end of the tensor release node, and the output end of the tensor release node is connected to the input end of the downstream computational logic node of the computational logic node that uses the tensor in the task flow to which the computational logic node generating the tensor belongs.

[0032] With the system according to the present disclosure, based on the tensor to be processed, by controlling the deployment of the tensor release node in the task flow with a faster processing speed that all involves the same tensor, the situation of a sharp increase in memory requirements caused by the faster task flow can be controlled. The processing method of the present disclosure automatically eliminates the need for manual adjustment of the memory explosion situation for the computational graph or the program corresponding to the computational graph. More precisely, it enables the memory requirement explosion avoidance process in the program to be automatically performed without manual intervention, greatly increasing the labor cost of program error correction and reducing the maintenance cost and the possibility of errors in the multi-flow parallel system. Through the system of the present disclosure, the computation and video memory management are uniformly regarded as instruction management and are video memory safe. Moreover, compared with other processing methods, the present disclosure has been dynamically processed without the need for developers to manually repeat the implementation. The maintenance cost and the possibility of errors are reduced.

[0033] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that for those of ordinary skill in the art, all or any steps or components of the method and apparatus of the present disclosure can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present disclosure.

[0034] Therefore, the object of the present disclosure can also be achieved by running a program or a set of programs on any computing device. The computing device may be a well-known general-purpose device. Therefore, the object of the present disclosure can also be achieved only by providing a program product containing program code for implementing the method or apparatus. That is to say, such a program product also constitutes the present disclosure, and a storage medium storing such a program product also constitutes the present disclosure. Obviously, the storage medium may be any well-known storage medium or any storage medium developed in the future.

[0035] It should also be noted that in the devices and methods of the present disclosure, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure. Moreover, the steps of performing the above series of processes can naturally be executed chronologically in the described order, but it is not necessary to be executed necessarily in chronological order. Certain steps can be executed in parallel or independently of each other.

[0036] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A system for avoiding a surge in memory requirements in multi-stream parallel data processing, comprising: The task flow computing engine, as the task is executed, when inserting the computing logic node task contained in the task into the corresponding task flow to which it belongs, confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, the computing logic node is directly inserted into the task flow to which it belongs; A tensor information acquisition component, when the task flow computing engine confirms that the computing logic node to be inserted is not a tensor generation logic node, acquires information of all computing logic nodes that need to use the generated tensor; as well as The tensor release node reconstruction component reconstructs the tensor release node based on the information obtained by the tensor information acquisition component, so that the output ends of the computing logic nodes that need to use the tensor on all task flows are connected to the input ends of the tensor release node, and the output end of the tensor release node is connected between the computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs and its downstream computing logic node, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow to which it belongs, so that the subsequent computing logic nodes in the task flow where the reconstructed tensor release node is located apply for memory execution after the tensor release node completes the tensor release. When the tensor is used by other downstream computing logic nodes of the computing logic node generated by the tensor, the task flow computing engine inserts the reconstructed tensor release node immediately before the computing logic node that last uses the tensor in the task flow to which it belongs.

2. The system for avoiding a surge in memory requirements in multi-stream parallel data processing according to claim 1, wherein: After the task flow computing engine has emitted a reconstructed tensor release node, when receiving a release node for the same tensor again, the task flow computing engine discards the tensor release node.

3. A method for avoiding a surge in memory requirements in multi-stream parallel data processing, comprising: As the task is executed, when the task flow computing engine inserts the computing logic node task contained in the task into the corresponding task flow to which it belongs, it confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, the computing logic node is directly inserted into the task flow to which it belongs; When the task flow computing engine confirms that the computing logic node to be inserted is not a tensor generation logic node, the information of all computing logic nodes that need to use the generated tensor is obtained through the tensor information acquisition component; as well as Based on the information obtained by the tensor information acquisition component, the tensor release node is reconstructed through the tensor release node reconstruction component, so that the output ends of the computing logic nodes that need to use the tensor on all task flows are connected to the input ends of the tensor release node, and the output end of the tensor release node is connected between the computing logic node that uses the tensor and its downstream computing logic node in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow to which it belongs, so that the subsequent computing logic nodes in the task flow where the reconstructed tensor release node is located apply for memory execution after the tensor release node completes the tensor release. When the tensor is used by other downstream computing logic nodes of the computing logic node generated by the tensor, the task flow computing engine inserts the reconstructed tensor release node immediately before the computing logic node that uses the tensor last in the task flow to which it belongs.

4. The method for avoiding a surge in memory requirements in multi-stream parallel data processing according to claim 3, wherein: After the task flow computing engine has emitted a reconstructed tensor release node, when receiving a release node for the same tensor again, the task flow computing engine discards the tensor release node.

5. A system for avoiding a surge in memory requirements in multi-stream parallel data processing, comprising: The task flow computing engine, as the task is executed, when inserting the computing logic node task contained in the task into the corresponding task flow to which it belongs, confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, the computing logic node is directly inserted into the task flow to which it belongs; A tensor information acquisition component, when the task flow computing engine confirms that the computing logic node to be inserted is not a tensor generation logic node, acquires information of all computing logic nodes that need to use the generated tensor; as well as A tensor release node reconstruction component reconstructs the tensor release node based on the information obtained by the tensor information acquisition component, so that the output ends of the computing logic nodes that need to use the tensor on all task flows are connected to the input ends of the tensor release node, and the output ends of the tensor release node are connected to the input ends of the downstream computing logic nodes of the computing logic nodes that use the tensor in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow containing the computing logic nodes that use the tensor, so that the computing logic nodes after the computing logic node that last uses the tensor in the task flow to which the computing logic node that generates the tensor belongs apply for memory after the tensor release node completes the tensor release.

6. A method for avoiding a surge in memory requirements in multi-stream parallel data processing, comprising: As the task is executed, when the task flow computing engine inserts the computing logic node task contained in the task into the corresponding task flow to which it belongs, it confirms whether the computing logic node to be inserted is a tensor generation logic node. If not, the computing logic node is directly inserted into the task flow to which it belongs; When the task flow computing engine confirms that the computing logic node to be inserted is a tensor generation logic node, the information of all computing logic nodes that need to use the generated tensor is obtained through the tensor information acquisition component; as well as Based on the information obtained by the tensor information acquisition component, the tensor release node is reconstructed through the tensor release node reconstruction component, so that the output ends of the computing logic nodes that need to use the tensor on all task flows are connected to the input end of the tensor release node, and the output end of the tensor release node is connected to the input end of the downstream computing logic node of the computing logic node that uses the tensor in the task flow to which the computing logic node that generates the tensor belongs, so that the task flow computing engine inserts the reconstructed tensor release node into the task flow containing the computing logic node that uses the tensor, so that the computing logic node after the computing logic node that last uses the tensor in the task flow to which the computing logic node that generates the tensor belongs applies for memory after the tensor release node completes the tensor release.

Citation Information

Patent Citations

  • Computer-readable recording medium for recording learning program and learning method

    CN112465105A

  • Synchronous deployment system and method for multi-stream parallelism

    CN114035810A