Model processing method and system for data stream architecture and related equipment

By constructing delay tables and computational graph segmentation, the problem of memory bandwidth limitation is solved, efficient pipeline parallel inference of deep learning models is realized, and the model inference efficiency is improved.

CN120471174APending Publication Date: 2025-08-12YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711343.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the process of deep learning model inference, memory bandwidth becomes a limiting factor, resulting in performance degradation, especially when processing large-scale data and complex models, which reduces the efficiency of model inference.

Method used

By building model nodes to execute delay tables and communication delay tables, determine the minimum latency requirements of model services, perform deep learning model calculation graph segmentation, and make full use of data flow architecture computing card resources through pipeline parallel inference.

Benefits of technology

The concurrency quantity and efficiency of model inference services are improved, and pipeline parallel inference is performed by calculating graph segmentation results and data flow architecture hardware characteristics, which improves the inference efficiency of deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471174A_ABST
    Figure CN120471174A_ABST
Patent Text Reader

Abstract

The invention discloses a model processing method and system for a data stream architecture and related equipment, and relates to the technical field of model reasoning, and the method comprises the steps: obtaining model parameters of a deep learning model, determining hardware characteristics of the data stream architecture, constructing a model node execution delay table and a communication delay table of the deep learning model, and determining the minimum delay requirement of model service. According to a preset segmentation mode, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table, deep learning model calculation graph segmentation is carried out, a calculation graph segmentation result is obtained, and pipeline parallel reasoning is carried out through the calculation graph segmentation result and the data flow architecture hardware characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model reasoning technology, and more specifically, to a model processing method, system and related equipment for a data flow architecture. Background Art

[0002] Deep learning model inference refers to the computational process of using a trained model to make predictions or decisions about new input data. Its core goal is to solve practical problems through efficient, accurate, and real-time output.

[0003] Currently, bandwidth is often a bottleneck limiting the performance of deep learning model inference. Therefore, dataflow architecture hardware has been developed to address the problem of intermediate model results needing to be saved to global memory and then loaded from global memory upon subsequent use. Dataflow architecture hardware can directly transfer the output data of one compute node in a model to the local memory of the next compute node, leveraging the high bandwidth of local memory to improve data transfer efficiency.

[0004] However, memory bandwidth can become a limiting factor when performing deep learning model inference on dataflow-based hardware. When processing large-scale data and complex models, memory requirements can grow rapidly, and dataflow-based hardware may not be able to provide sufficient memory bandwidth to meet the demand, resulting in performance degradation and, consequently, reduced efficiency in deep learning model inference.

[0005] Therefore, how to improve the efficiency of deep learning model reasoning is an urgent problem that needs to be solved in this application. Summary of the Invention

[0006] In view of this, the present application discloses a model processing method, system and related equipment of a data flow architecture, aiming to improve the efficiency of deep learning model reasoning.

[0007] In order to achieve the above purpose, the disclosed technical solutions are as follows:

[0008] In a first aspect, the present application discloses a model processing method for a data stream architecture, the method comprising:

[0009] Get the model parameters of the deep learning model;

[0010] Determine data flow architecture hardware characteristics;

[0011] Construct model node execution delay table and communication delay table of deep learning model;

[0012] Determine the minimum latency requirements for model serving;

[0013] Segment the deep learning model calculation graph according to the preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table to obtain a calculation graph segmentation result;

[0014] Pipeline parallel reasoning is performed based on the computational graph segmentation results and the hardware characteristics of the data flow architecture.

[0015] Preferably, determining the data flow architecture hardware characteristics includes:

[0016] Determine a data flow architecture computing card; wherein the data flow architecture computing card includes multiple computing cores and dynamic random access memory; each computing core includes a local static random access memory;

[0017] According to the data flow architecture computing card, data flow architecture hardware characteristics are determined.

[0018] Preferably, the model node execution delay table and communication delay table for constructing the deep learning model include:

[0019] Analyze each model node in a deep learning model to obtain execution latency under different numbers of computing cores;

[0020] Constructing a model node execution delay table of the deep learning model according to the execution delay times under the different numbers of computing cores;

[0021] Analyze the communication delay caused by card-to-card communication between different model nodes;

[0022] The communication delay table of the deep learning model is constructed through the communication delay caused by communication between cards.

[0023] Preferably, determining the minimum delay requirement of the model service includes:

[0024] Obtain user application scenarios and user needs;

[0025] Based on the user application scenario and the user needs, determine the minimum latency requirement for the model service of the deep learning model on the data flow architecture hardware.

[0026] Preferably, the deep learning model calculation graph is segmented according to the preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table to obtain the calculation graph segmentation result, including:

[0027] Look up the execution delay table of the model node to obtain the execution time of the node;

[0028] Get the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set and the number of nodes;

[0029] Look up the communication delay table to obtain the communication delay;

[0030] An integer linear programming solution is performed based on the number of computing cores of each data flow architecture computing card, the number of actually available data flow architecture computing cards, the deep learning model node set, the number of nodes, the communication delay, and the execution time of the node to complete the segmentation of the deep learning model computing graph and obtain a computing graph segmentation result; wherein the computing graph segmentation result represents the computing card where each node is located, the number of computing cores occupied by each node, and the segmentation method between different nodes.

[0031] Preferably, the pipeline parallel reasoning is performed using the computation graph segmentation result and the data flow architecture computing card, including:

[0032] Determine the computing nodes after segmentation based on the segmentation result of the computing graph;

[0033] Mapping the split computing nodes to the data flow architecture computing cards;

[0034] Pipeline parallel reasoning is performed through pipeline parallelism and mapped data flow architecture computing cards.

[0035] A second aspect of the present application discloses a model processing system for a data flow architecture, the system comprising:

[0036] An acquisition unit, used to obtain model parameters of a deep learning model;

[0037] A first determining unit, configured to determine hardware characteristics of a data flow architecture;

[0038] A construction unit for constructing a model node execution delay table and a communication delay table of a deep learning model;

[0039] A second determining unit is used to determine the minimum latency requirement of the model service;

[0040] A segmentation unit is used to segment the deep learning model calculation graph according to a preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table, and obtain a calculation graph segmentation result;

[0041] A parallel reasoning unit is used to perform pipeline parallel reasoning based on the computation graph segmentation results and the hardware characteristics of the data flow architecture.

[0042] Preferably, the first determining unit includes:

[0043] A first determining module is configured to determine a data flow architecture computing card, wherein the data flow architecture computing card includes a plurality of computing cores and dynamic random access memories, and each computing core includes a local static random access memory;

[0044] The second determining module is used to determine the data flow architecture hardware characteristics according to the data flow architecture computing card.

[0045] A third aspect of the present application discloses a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the model processing method of the data flow architecture as described in any one of the first aspects.

[0046] The fourth aspect of the present application discloses an electronic device comprising a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform a model processing method of a data flow architecture as described in any one of the first aspects.

[0047] Through the above technical solution, it can be seen that the present application discloses a model processing method, system and related equipment of a data flow architecture, obtains the model parameters of the deep learning model, determines the hardware characteristics of the data flow architecture, constructs the model node execution delay table and communication delay table of the deep learning model, determines the minimum delay requirements of the model service, and performs deep learning model calculation graph segmentation according to the preset segmentation method, model parameters, minimum delay requirements of the model service, model node execution delay table and communication delay table to obtain the calculation graph segmentation results, and performs pipeline parallel reasoning based on the calculation graph segmentation results and the hardware characteristics of the data flow architecture.

[0048] Through the above scheme, the minimum delay requirement of the model service is determined by the model node execution delay table and communication delay table of the deep learning model. When the data flow architecture hardware characteristics, deep learning model and deep learning model inference delay requirements are determined, the calculation graph of the deep learning model is quickly segmented, and the segmented computing nodes are mapped to the given data flow architecture computing card. Through the pipeline parallel method, the data flow architecture computing card resources are fully utilized, the number of concurrent model inference services is increased, and pipeline parallel inference is performed based on the calculation graph segmentation results and the data flow architecture hardware characteristics, thereby improving the efficiency of deep learning model inference. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0050] Figure 1 A flow chart of a model processing method for a data stream architecture disclosed in an embodiment of the present application;

[0051] Figure 2 This is an example diagram of the computational graph segmentation disclosed in the embodiments of this application;

[0052] Figure 3 A schematic diagram of pipeline parallel reasoning disclosed in an embodiment of the present application;

[0053] Figure 4 A schematic diagram of the structure of a model processing system of a data flow architecture disclosed in an embodiment of the present application;

[0054] Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0056] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0057] As can be seen from the background technology, memory bandwidth can become a limiting factor when performing deep learning model inference using dataflow-based hardware. When processing large amounts of data and complex models, memory requirements can grow rapidly, and dataflow-based hardware may not be able to provide sufficient memory bandwidth to meet these requirements, resulting in performance degradation and, consequently, reduced efficiency in deep learning model inference. Therefore, improving the efficiency of deep learning model inference is an urgent issue that needs to be addressed in this application.

[0058] In order to solve the above problems, the present application discloses a model processing method, system and related equipment of a data flow architecture, which determines the minimum delay requirement of the model service through the model node execution delay table and communication delay table of the deep learning model, and quickly performs the calculation graph segmentation of the deep learning model under the condition of determining the data flow architecture hardware characteristics, the deep learning model and the deep learning model reasoning delay requirements, and maps the segmented calculation nodes to the given data flow architecture calculation card. By using the pipeline parallel method, the data flow architecture calculation card resources are fully utilized, the number of concurrent model reasoning services is increased, and pipeline parallel reasoning is performed based on the calculation graph segmentation results and the data flow architecture hardware characteristics, thereby improving the efficiency of deep learning model reasoning. The specific implementation method is specifically described through the following embodiments.

[0059] It should be noted that the model processing method, system and related equipment of a data flow architecture provided in this application can be used in technical fields such as model reasoning, pipeline parallelism, and data flow architecture. The above is only an example and does not limit the application field of the model processing method, system and related equipment of a data flow architecture provided in this application.

[0060] refer to Figure 1 FIG. 1 is a flow chart of a model processing method for a data flow architecture disclosed in an embodiment of the present application. The model processing method for the data flow architecture mainly includes the following steps:

[0061] S101: Obtain model parameters of a deep learning model.

[0062] Among them, model parameters include the network architecture, number of nodes, operator type, etc. of the deep learning model.

[0063] The types of network architectures of deep learning models include Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Generative Adversarial Networks (GAN), etc.

[0064] Operator types include linear, convolution, nonlinear, etc.

[0065] S102: Determine data flow architecture hardware characteristics.

[0066] In S102, a data flow architecture computing card is determined, and data flow architecture hardware characteristics are determined based on the data flow architecture computing card.

[0067] This application applies to the following specific data flow architecture hardware, which has the following characteristics:

[0068] A data flow architecture computing card includes multiple computing cores and dynamic random access memory (DRAM); each computing core includes local static random access memory (SRAM). The computing cores can perform asynchronous calculations, increasing the degree of parallelism. In terms of the memory system, a computing card has a large DRAM. Data can be transferred between DRAM and SRAM, and data can be transferred between different SRAMs. Furthermore, this application supports communication between multiple computing cards, and this application does not specifically limit the communication method.

[0069] S103: Construct a model node execution delay table and a communication delay table of the deep learning model.

[0070] To fully leverage the hardware advantages of the dataflow architecture, after a model node outputs data, it bypasses DRAM and is instead directly transmitted to the local SRAM of the model node that requires it as input. Using this dataflow approach, actual testing analyzes the execution latency of each model node under different numbers of compute cores.

[0071] In addition, this application supports multiple computing cards in parallel, involving P2P communication between cards, and the physical communication method between computing cards is not limited. Similarly, through actual testing, the communication delay caused by card-to-card communication between different nodes is analyzed.

[0072] The specific process of constructing the model node execution delay table and communication delay table of the deep learning model is shown in A1-A4.

[0073] A1: Analyze each model node in the deep learning model to obtain the execution latency under different numbers of computing cores.

[0074] In A1, each model node in the deep learning model can be analyzed through actual testing to measure the execution latency under different numbers of computing cores.

[0075] A2: Build a model node execution delay table for the deep learning model based on the execution delay time under different numbers of computing cores.

[0076] The functions of the model node execution delay table are as follows:

[0077] (1) Performance analysis and optimization: By recording the delay time of each node processing tasks, we can intuitively understand the performance of each node, thereby discovering performance bottlenecks, making targeted optimizations, and improving the operating efficiency of the system.

[0078] (2) Resource Scheduling and Management: Allocate tasks and resources based on the node execution delay table to ensure that tasks can be completed in the shortest possible time. For nodes with higher latency, the amount of tasks assigned can be reduced, or they can be upgraded or optimized.

[0079] (3) Fault diagnosis and troubleshooting: When a system has performance issues or a fault, the node execution delay table can help quickly locate the problem. By comparing the delay data under normal conditions and during a fault, the abnormal node can be found, and in-depth troubleshooting and repair can be carried out.

[0080] A3: Analyze the communication delay caused by card-to-card communication between different model nodes.

[0081] In A3, actual testing can be used to analyze the communication delay caused by card-to-card communication between different model nodes.

[0082] A4: Build a communication delay table for the deep learning model based on the communication delay caused by communication between cards.

[0083] The communication delay table for deep learning models serves the following purposes:

[0084] (1) Network performance evaluation: The communication delay table records the communication delay between different nodes and can reflect the overall performance of the network. By analyzing the communication delay data, it can be used to evaluate whether the network meets business needs and whether network upgrades or optimizations are needed.

[0085] (2) Communication protocol selection and optimization: Different communication protocols exhibit different communication delay characteristics in different network environments. By comparing the communication delays of different protocols, the most suitable communication protocol can be selected, or existing protocols can be optimized to reduce communication delay.

[0086] (3) Distributed system design and optimization: In a distributed system, communication delay between nodes is a significant factor affecting system performance. Using a communication delay table, we can optimize the system architecture and communication methods, reduce unnecessary communication overhead, and improve system response speed and throughput.

[0087] S104: Determine the minimum latency requirement for the model service.

[0088] In S104, user application scenarios and user needs are obtained, and based on the user application scenarios and user needs, the minimum latency requirements for the model service of the deep learning model on the data flow architecture hardware are determined.

[0089] The minimum latency requirement refers to the minimum time it takes for a deep learning model to complete a request.

[0090] S105: Segment the deep learning model calculation graph according to the preset segmentation method, model parameters and minimum latency requirements of the model service to obtain the calculation graph segmentation results.

[0091] The default partitioning method uses an integer linear programming method to determine the compute card on which each node resides, the number of compute cores each node occupies, and the partitioning method between different nodes. The specific process of using the integer linear programming method to obtain the computational graph partitioning results is shown in Figures B1-B3.

[0092] B1: Look up the model node execution delay table to obtain the node execution time.

[0093] The execution time of the node is expressed as express.

[0094] B2: Obtain the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set, and the number of nodes.

[0095] In the process of splitting the deep learning model calculation graph, assuming that the minimum delay requirement is T, the number of computing cores on each computing card is Indicates the actual number of available computing cards. Represents; deep learning model node collection is used Indicates the number of nodes. The execution time of a node is expressed as Represents, where i represents the i-th node, b represents the batch size, and m represents the number of computing cores used; This can be obtained by querying the model node execution delay table; The value is 0 or 1. If A value of 1 indicates that node i uses a batch size of b and m number of computing cores; if A value of 0 indicates that the batch size used by node i is not b or the number of computing cores used is not m.

[0096] B3: Look up the communication delay table to obtain the communication delay.

[0097] B4: Based on the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set, the number of nodes, the communication delay, and the node execution time, perform integer linear programming to complete the segmentation of the deep learning model computing graph and obtain the computing graph segmentation result; among them, the computing graph segmentation result indicates the computing card where each node is located, the number of computing cores occupied by each node, and the segmentation method between different nodes.

[0098] The computation graph segmentation result indicates the computing card where each node is located, the number of computing cores occupied by each node, and the segmentation method between different nodes.

[0099] The above problem is an integer linear programming problem, which can be solved using the integer linear programming method to obtain the computing card where each node is located, the number of computing cores occupied by each node, and the partitioning method between different nodes.

[0100] The specific formula solved by the integer linear programming method is shown in formula (1).

[0101] (1)

[0102] Assume that the number of computing cards occupied by a single model inference is C; the number of computing cards is ;use represents the communication delay between node i and node j, which can be obtained by querying the communication delay table; Indicates whether nodes i and j are on the same computing card. The value is 0 or 1. 0 means they are on the same computing card, and 1 means they are on different computing cards. If the two nodes are on the same computing card, their communication delay is already included in the node execution time. Indicates the number of concurrent connections supported; Indicates the total number of computing cores occupied by the node on computing card c.

[0103] The calculation graph segmentation results are as follows Figure 2 shown.

[0104] Figure 2 In the example, op1, op2, op3, op4, op5, and op6 are all nodes.

[0105] The six nodes op1, op2, op3, op4, op5, and op6 are split across two data flow architecture computing cards.

[0106] S106: Perform pipeline parallel reasoning based on the computational graph segmentation results and the data flow architecture hardware characteristics.

[0107] Specifically, the process of pipeline parallel reasoning is performed by computing the graph segmentation results and the hardware characteristics of the data flow architecture, as shown in C1-C3.

[0108] C1: Determine the computing nodes after segmentation based on the computational graph segmentation results.

[0109] C2: Map the split computing nodes to the data flow architecture computing cards.

[0110] C3: Performs pipeline parallel reasoning through pipeline parallelism and mapped data flow architecture computing cards.

[0111] According to the above calculation graph segmentation results, pipeline parallel reasoning is performed. Figure 3 Example shown.

[0112] Figure 3 Here, b1, b2, b3, b4, b5, b6, b7, and b8 represent the compute nodes after segmentation.

[0113] Figure 3 In the example, node op1 executes the first batch request first. After the execution is completed, node op2 starts to execute the first batch request. At the same time, op1 executes the second batch request. This process is repeated in sequence to implement pipeline parallel reasoning of different batches on the data flow architecture computing card. At the same time, the characteristics of the data flow architecture can be utilized to avoid the transmission of intermediate results of different nodes in DRAM and SRAM.

[0114] This application rapidly segments the computational graph of a deep learning model, given the hardware, model, and model inference latency requirements, and maps the segmented computational nodes to a given dataflow architecture compute card. By leveraging pipeline parallelism, the dataflow architecture compute card resources are fully utilized, increasing the concurrency of model inference services. This method can be applied to deep learning tasks such as image processing, natural language processing, and artificial intelligence generated content (AIGC), demonstrating both practical and innovative value.

[0115] The beneficial effects of the embodiments of the present application are as follows: the minimum delay requirement of the model service is determined by the model node execution delay table and the communication delay table of the deep learning model; while determining the data flow architecture hardware characteristics, the deep learning model and the deep learning model reasoning delay requirements, the calculation graph of the deep learning model is quickly segmented, and the segmented calculation nodes are mapped to the given data flow architecture computing card; through the pipeline parallel method, the data flow architecture computing card resources are fully utilized, the number of concurrent model reasoning services is increased, and pipeline parallel reasoning is performed based on the calculation graph segmentation results and the data flow architecture hardware characteristics, thereby improving the efficiency of deep learning model reasoning.

[0116] Based on the above embodiment Figure 1The disclosed model processing method of a data flow architecture, the embodiment of the present application also discloses a model processing system of a data flow architecture, such as Figure 4 As shown, the model processing system of the data flow architecture includes:

[0117] An acquisition unit 401 is used to acquire model parameters of a deep learning model;

[0118] A first determining unit 402 is configured to determine hardware characteristics of a data flow architecture;

[0119] A construction unit 403 is used to construct a model node execution delay table and a communication delay table of the deep learning model;

[0120] A second determining unit 404 is configured to determine a minimum delay requirement for the model service;

[0121] A segmentation unit 405 is configured to segment the deep learning model computation graph according to a preset segmentation method, model parameters, minimum delay requirements of the model service, a model node execution delay table, and a communication delay table, and obtain a computation graph segmentation result;

[0122] The parallel reasoning unit 406 is used to perform pipeline parallel reasoning based on the computation graph segmentation results and the data flow architecture hardware characteristics.

[0123] Furthermore, the first determining unit 402 includes:

[0124] A first determining module is configured to determine a data flow architecture computing card, wherein the data flow architecture computing card includes a plurality of computing cores and dynamic random access memories, and each computing core includes a local static random access memory;

[0125] The second determining module is used to determine the data flow architecture hardware characteristics according to the data flow architecture computing card.

[0126] Furthermore, the construction unit 403 includes:

[0127] The first analysis module is used to analyze each model node in the deep learning model to obtain the execution delay time under different numbers of computing cores;

[0128] A first construction module is used to construct a model node execution delay table of a deep learning model according to execution delay times under different numbers of computing cores;

[0129] The second analysis module is used to analyze the communication delay caused by card-to-card communication between different model nodes;

[0130] The second building module is used to build a communication delay table of the deep learning model through the communication delay caused by communication between cards.

[0131] Furthermore, the second determining unit 404 includes:

[0132] The first acquisition module is used to obtain user application scenarios and user needs;

[0133] The third determination module is used to determine the minimum delay requirements of the model service of the deep learning model on the data flow architecture hardware based on the user application scenario, user needs, model node execution delay table and communication delay table.

[0134] Furthermore, the segmentation unit 405 includes:

[0135] The second acquisition module is used to look up the model node execution delay table to obtain the node execution time;

[0136] The third acquisition module is used to obtain the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set and the number of nodes;

[0137] A fourth acquisition module is used to look up the communication delay table to obtain the communication delay;

[0138] The solution module is used to perform integer linear programming solution based on the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set, the number of nodes, the communication delay and the execution time of the node, so as to complete the segmentation of the deep learning model computing graph and obtain the computing graph segmentation result; wherein, the computing graph segmentation result indicates the computing card where each node is located, the number of computing cores occupied by each node, and the segmentation method between different nodes.

[0139] Furthermore, the parallel reasoning unit 406 includes:

[0140] A fourth determination module is used to determine the computation nodes after segmentation based on the computation graph segmentation result;

[0141] A mapping module is used to map the split computing nodes to the data flow architecture computing cards;

[0142] The parallel reasoning module is used to perform pipeline parallel reasoning through pipeline parallelism and mapped data flow architecture computing cards.

[0143] The beneficial effects of the embodiments of the present application are as follows: the minimum delay requirement of the model service is determined by the model node execution delay table and the communication delay table of the deep learning model; while determining the data flow architecture hardware characteristics, the deep learning model and the deep learning model reasoning delay requirements, the calculation graph of the deep learning model is quickly segmented, and the segmented computing nodes are mapped to the given data flow architecture computing card; through the pipeline parallel method, the computing card resources are fully utilized, the number of concurrent model reasoning services is increased, and pipeline parallel reasoning is performed based on the calculation graph segmentation results and the data flow architecture hardware characteristics, thereby improving the efficiency of deep learning model reasoning.

[0144] An embodiment of the present application further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the model processing method of the data flow architecture as described above.

[0145] The present application also provides an electronic device, the structure of which is shown in FIG. Figure 5 As shown, it specifically includes a memory 501 and one or more instructions 502, wherein the one or more instructions 502 are stored in the memory 501 and are configured to be executed by one or more processors 503 to execute the one or more instructions 502 to perform the model processing method of the above-mentioned data flow architecture.

[0146] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0147] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.

[0148] The steps in the methods of the various embodiments of the present application can be adjusted in sequence, combined, and deleted according to actual needs.

[0149] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0150] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

[0151] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data flow architecture model processing method, characterized in that: The method comprises: Get the model parameters of the deep learning model; Determine data flow architecture hardware characteristics; Construct model node execution delay table and communication delay table of deep learning model; Determine the minimum latency requirements for model serving; Segment the deep learning model calculation graph according to the preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table to obtain a calculation graph segmentation result; Pipeline parallel reasoning is performed based on the computational graph segmentation results and the hardware characteristics of the data flow architecture.

2. The method according to claim 1, characterized in that Determining the data flow architecture hardware characteristics includes: Determine a data flow architecture computing card; wherein the data flow architecture computing card includes multiple computing cores and dynamic random access memory; each computing core includes a local static random access memory; According to the data flow architecture computing card, data flow architecture hardware characteristics are determined.

3. The method according to claim 1, characterized in that The model node execution delay table and communication delay table for constructing the deep learning model include: Analyze each model node in a deep learning model to obtain execution latency under different numbers of computing cores; Constructing a model node execution delay table of the deep learning model according to the execution delay times under the different numbers of computing cores; Analyze the communication delay caused by card-to-card communication between different model nodes; The communication delay table of the deep learning model is constructed through the communication delay caused by communication between cards.

4. The method according to claim 1, wherein The minimum latency requirements for determining the model service include: Obtain user application scenarios and user needs; Based on the user application scenario and the user needs, determine the minimum latency requirement for the model service of the deep learning model on the data flow architecture hardware.

5. The method according to claim 1, wherein The deep learning model calculation graph is segmented according to the preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table to obtain the calculation graph segmentation result, including: Look up the execution delay table of the model node to obtain the execution time of the node; Get the number of computing cores of each data flow architecture computing card, the actual number of available data flow architecture computing cards, the deep learning model node set and the number of nodes; Look up the communication delay table to obtain the communication delay; An integer linear programming solution is performed based on the number of computing cores of each data flow architecture computing card, the number of actually available data flow architecture computing cards, the deep learning model node set, the number of nodes, the communication delay, and the execution time of the node to complete the segmentation of the deep learning model computing graph and obtain a computing graph segmentation result; wherein the computing graph segmentation result represents the computing card where each node is located, the number of computing cores occupied by each node, and the segmentation method between different nodes.

6. The method according to claim 1, characterized in that The pipeline parallel reasoning is performed using the computation graph segmentation result and the data flow architecture computing card, including: Determine the computing nodes after segmentation based on the segmentation result of the computing graph; Mapping the split computing nodes to the data flow architecture computing cards; Pipeline parallel reasoning is performed through pipeline parallelism and mapped data flow architecture computing cards.

7. A data flow architecture model processing system, characterized in that: The system comprises: An acquisition unit, used to obtain model parameters of a deep learning model; A first determining unit, configured to determine hardware characteristics of a data flow architecture; A construction unit for constructing a model node execution delay table and a communication delay table of a deep learning model; A second determining unit is used to determine the minimum latency requirement of the model service; A segmentation unit is used to segment the deep learning model calculation graph according to a preset segmentation method, the model parameters, the minimum delay requirement of the model service, the model node execution delay table and the communication delay table, and obtain a calculation graph segmentation result; A parallel reasoning unit is used to perform pipeline parallel reasoning based on the computation graph segmentation results and the hardware characteristics of the data flow architecture.

8. The system according to claim 7, characterized in that The first determining unit includes: A first determining module is configured to determine a data flow architecture computing card, wherein the data flow architecture computing card includes a plurality of computing cores and dynamic random access memories, and each computing core includes a local static random access memory; The second determining module is used to determine the data flow architecture hardware characteristics according to the data flow architecture computing card.

9. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the model processing method of the data flow architecture according to any one of claims 1 to 6.

10. An electronic device, characterized in that: It includes a memory and one or more instructions, wherein the one or more instructions are stored in the memory and are configured to be executed by one or more processors to perform the model processing method of the data flow architecture as described in any one of claims 1 to 6.