Method, device, computer device and storage medium for offline model inference

By employing a parallel and pipelined design with a multi-stage pipeline architecture, the inefficiency of NPU chips in processing complex deep learning network models is resolved, achieving more efficient model inference computation.

CN115511079BActive Publication Date: 2026-02-27CAMBRIAN (KUNSHAN) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110629502.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-07
Publication Date
2026-02-27
Estimated Expiration
2041-08-12

AI Technical Summary

Technical Problem

As deep learning network models become more complex, the inference efficiency of NPU chips gradually decreases, making it difficult to efficiently process video and image data.

Method used

A multi-stage pipeline architecture is adopted, including multiple first and second pipelines. Through parallel processing and cascading design, the parallel and pipelined operation of multiple pipelines is realized, thereby improving the efficiency of model inference.

Benefits of technology

It greatly improves the speed and efficiency of model inference and optimizes the computational performance of deep learning network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115511079B_ABST
    Figure CN115511079B_ABST
Patent Text Reader

Abstract

The application relates to a method, device, computer equipment and storage medium for offline model reasoning. The method comprises the following steps: obtaining to-be-processed data of multiple offline models; performing first parallel processing on the to-be-processed data of the multiple offline models on multiple first pipelines; storing processing results output by the first pipelines into a first cache queue; reading the processing results output by the first pipelines from the first cache queue on multiple second pipelines to perform second parallel processing, so as to complete reasoning of the multiple offline models. The method for offline model reasoning realizes a reasoning calculation method of multiple pipeline parallelization and pipelining operation in an offline model reasoning application scenario. Compared with a method of single-pipeline execution of reasoning calculation, the method greatly improves the speed and efficiency of model reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, in particular to an offline model reasoning method and device, computer equipment and a storage medium. BACKGROUND

[0002] With the continuous development of machine learning algorithms, more and more architectures of machine learning chips have gradually emerged. Among them, the embedded neural network processor (NPU) is widely used in the recognition and processing of video data and massive images due to its adoption of the "data-driven parallel computing" architecture.

[0003] At present, in the process of processing video data or image data using an NPU chip, the inference of a deep learning network model is usually involved, that is, all calculations in the deep learning network model are mapped to the NPU chip for operation, and the operation result is the processing result of the video data or image data.

[0004] However, with the increasing complexity of deep learning network models, the inference efficiency of NPU chips for deep learning network models is also increasingly low. SUMMARY

[0005] Therefore, it is necessary to provide an offline model reasoning method, device, computer equipment and storage medium capable of improving the inference efficiency of a deep learning network model to solve the above technical problems.

[0006] In a first aspect, an offline model reasoning method is provided, which adopts a multi-stage pipeline including a plurality of first pipelines and a plurality of second pipelines, and the method comprises:

[0007] obtaining to-be-processed data of a plurality of offline models;

[0008] performing first parallel processing on the to-be-processed data of the plurality of offline models on the plurality of first pipelines, and storing the processing results output by each of the first pipelines to a first cache queue;

[0009] performing second parallel processing on the processing results output by each of the first pipelines read from the first cache queue on the plurality of second pipelines to complete the inference of the plurality of models.

[0010] In a second aspect, an offline model reasoning device is provided, which comprises:

[0011] an obtaining module configured to obtain to-be-processed data of a plurality of offline models;

[0012] The first parallel processing module is configured to perform first parallel processing on the to-be-processed data of the plurality of offline models on a plurality of first pipelines, and store the processing results output by each of the first pipelines into a first cache queue.

[0013] The second parallel processing module is configured to perform second parallel processing on the processing results output by each of the first pipelines by reading the processing results from the first cache queue on a plurality of second pipelines, so as to complete the inference of the plurality of offline models.

[0014] In a third aspect, a computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method in the first aspect when executing the computer program.

[0015] In a fourth aspect, a computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method in the first aspect.

[0016] The method, device, computer device and storage medium for offline model inference provided in the present application realize parallel processing of data by setting a plurality of first pipelines and a plurality of second pipelines, and the second pipelines and the first pipelines are cascaded, that is, the second pipelines process the processing results output by the first pipelines, so as to realize the pipelining of data in different pipelines. Based on this, the method for offline model inference provided in the present application realizes a plurality of pipelining and pipelining inference calculation methods in the application scenario of offline model inference. Compared with the method for performing inference calculation by a single pipeline, the speed and efficiency of model inference are greatly improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1A FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0018] Figure 1B FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0019] Figure 2 FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0020] Figure 3 FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0021] Figure 4 FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0022] Figure 5 FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application; Figure 2 FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;

[0023] FIG. 1 is an internal structure diagram of a combination processing device in an embodiment of the present application;Figure 6 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0024] Figure 7 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0025] Figure 8 A schematic diagram of a method of offline model inference in an embodiment of the present application;

[0026] Figure 9 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0027] Figure 10 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0028] Figure 11 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0029] Figure 12 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0030] Figure 13 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0031] Figure 14 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0032] Figure 15 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0033] Figure 16 A schematic diagram of an offline model inference structure in an embodiment of the present application;

[0034] Figure 17 A schematic diagram of an offline model inference structure in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of, rather than all of, the embodiments of the present application. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0036] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0037] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0038] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0039] Figure 1A This is a structural diagram illustrating a combined processing apparatus 1200 according to an embodiment of this disclosure. Figure 1A As shown, the combined processing device 1200 includes a computing processing device 1202, an interface device 1204, other processing devices 1206, and a storage device 1208. Depending on the application scenario, the computing processing device may include one or more computing devices 1210, which can be configured to perform the functions described herein. Figures 2-14 The described operation.

[0040] The computing processing device 1202 is configured to perform user-specified operations, such as processes for performing offline model inference. In exemplary applications, the computing processing device 1202 is implemented primarily as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, the one or more computing devices 1210 included within the computing processing device 1202 can be implemented as artificial intelligence processor cores or portions of hardware structures of artificial intelligence processor cores. When multiple computing devices 1210 are implemented as artificial intelligence processor cores or portions of hardware structures of artificial intelligence processor cores, they can be considered to have a single-core structure or a homogeneous multi-core structure with respect to the computing processing device of the present disclosure.

[0041] The computing processing device 1202 interacts with other processing devices 1206 through the interface device 1204 to collectively perform user-specified operations. Depending on the implementation, the other processing devices 1206 of the present disclosure include one or more types of processors from among general-purpose and / or special-purpose processors such as central processing units (CPUs), graphics processing units (GPUs), artificial intelligence processors, and the like. These processors include, but are not limited to, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, and the like, and their number can be determined as needed. As previously mentioned, the computing processing device 1202 can be considered to have a single-core structure or a homogeneous multi-core structure with respect to the present disclosure. However, when the computing processing device 1202 and the other processing devices 1206 are considered collectively, they can be considered to form a heterogeneous multi-core structure.

[0042] The other processing devices 1206 can serve as an interface for external data and control for the computing processing device 1202 of the present disclosure, which can be embodied as a relevant computing device for artificial intelligence such as neural network operations, and perform basic controls including, but not limited to, data transfer, turning on and / or off of the computing devices 1210, and the like. In further embodiments, the other processing devices 1206 can also cooperate with the computing processing device 1202 to collectively perform computing tasks.

[0043] The interface device 1204 is configured to transmit data and control instructions between the computing processing device 1202 and the other processing device 1206. For example, the computing processing device 1202 can obtain input data from the other processing device 1206 via the interface device 1204 and write the input data into the storage device 1208 (or memory) on the computing processing device 1202. Further, the computing processing device 1202 can obtain control instructions from the other processing device 1206 via the interface device 1204 and write the control instructions into the control buffer on the computing processing device 1202. Alternatively or additionally, the interface device 1204 can also read data from the storage device of the computing processing device 1202 and transmit the data to the other processing device 1206.

[0044] The storage device 1208 is connected to the computing processing device 1202 and the other processing device 1206, respectively. The storage device 1208 is configured to store data of the computing processing device 1202 and / or the other processing device 1206. For example, the data can be data that cannot be stored in the internal or on-chip storage device of the computing processing device 1202 or the other processing device 1206.

[0045] In some embodiments, the disclosure also discloses a chip (e.g., the chip 1302 shown in Figure 1B FIG. 1). In one implementation, the chip is a System on Chip (SoC) and integrates one or more combined processing devices as shown in Figure 1A FIG. 1. The chip can be connected to other related components through an external interface device (e.g., the external interface device 1306 shown in Figure 1B FIG. 1). The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card, or a WiFi interface. In some application scenarios, other processing units (e.g., a video codec) and / or interface modules (e.g., a DRAM interface) can be integrated on the chip. In some embodiments, the disclosure also discloses a chip package structure including the chip described above. In some embodiments, the disclosure also discloses a board card including the chip package structure described above. The board card will be described in detail below. Figure 1B with reference to

[0046] Figure 1B FIG. 1 is a structural schematic diagram showing a board card 1300 according to an embodiment of the disclosure. As shown in Figure 1BAs shown, the board 1300 includes a storage device 1304 for storing data, which includes one or more storage cells 1310. The storage device 1304 can connect and transmit data with the controller 1308 and the aforementioned chip 1302 via, for example, a bus. Furthermore, the board also includes an external interface device 1306, configured for data relay or switching between the chip (or a chip in a chip package) and an external device 1312 (e.g., a server or computer). For example, data to be processed can be transmitted from the external device 1312 to the chip via the external interface device 1306. Alternatively, the calculation results of the chip can be transmitted back to the external device 1312 via the external interface device 1306. Depending on the application scenario, the external interface device 1306 can have different interface forms, such as a standard PCIe interface.

[0047] In one or more embodiments, the controller in the disclosed board can be configured to regulate the state of the chip. Therefore, in one application scenario, the controller includes a microcontroller (MCU) for regulating the operating state of the chip.

[0048] In one embodiment, such as Figure 2 As shown, an offline model inference method is provided. This method employs a multi-stage pipeline, which includes multiple first pipelines and multiple second pipelines. The method is applied to the processor in the board 1300 or the combination device 1200 in Figure 1, and includes the following steps:

[0049] S101: Obtain the data to be processed from multiple offline models.

[0050] Offline models include machine learning models, neural network models, or computational models built using other algorithms, such as clustering algorithms and search algorithms. The data to be processed, i.e., the data to be inferred or computed, includes different types of data such as images, text, video, and audio. In this embodiment, when the processor performs inference computations on different models, it can simultaneously acquire the data to be processed from multiple offline models, or acquire the data to be processed from multiple offline models sequentially. The type of the offline model indicates the specific type of operation performed on the input data. For example, when the offline model is a neural network model, the type can be any one of convolution, pooling, or concatenation operations. When the offline model is a computational model built using a clustering algorithm, the type can be one of accumulation and summation operations, averaging, or other operations. Multiple offline models can have the same or different types, and the types of the data to be processed from multiple offline models can also be the same or different.

[0051] S102, respectively on a plurality of first pipelines, a plurality of offline models of the data to be processed for the first parallel processing, and the first cache queue output of each first pipeline processing results stored.

[0052] Wherein, the pipeline is also called Pipeline. The first pipeline in the embodiment refers to the Pipeline for processing the data to be processed of the offline model. The data to be processed of each model is processed on each first pipeline, which includes preprocessing, inference calculation, post-processing and other work in the processing process, wherein, the preprocessing refers to a series of preprocessing and other processing work such as unit conversion, data resolution uniformity, noise removal, etc.; inference calculation refers to loading offline model to process data for model inference calculation; post-processing refers to a series of processing work such as drawing, statistical analysis and other processing work of the result of offline model inference. The first pipeline and the model are in one-to-one correspondence, that is, each first pipeline is used to process the data to be processed of the corresponding model, and the types of the plurality of first pipelines can be the same or different. The type of the first pipeline can be determined by the type of the model, that is, the type of the corresponding first pipeline is different when the type of the model is different.

[0053] In the embodiment, when the processor obtains the data to be processed of the plurality of offline models, the data to be processed of the plurality of offline models is processed in parallel on the plurality of first pipelines, so that the data to be processed of the corresponding offline model is processed on each first pipeline, each first pipeline performs parallel processing and does not affect each other, and the parallel processing result is obtained, that is, the processing result output by each first pipeline. Specifically, the processor first creates a first cache queue, and after obtaining the processing result output by each first pipeline, the processing result output by each first pipeline is directly stored in the first cache queue. Optionally, when the processor obtains the processing result output by each first pipeline at the same time, that is, each first pipeline completes the data processing at the same time, the processor stores the processing result output by each first pipeline in the first cache queue; when the processor obtains the processing result output by each first pipeline in sequence, the processor can also store the processing result output by each first pipeline in the first cache queue in sequence according to the order of the first pipeline completing the data processing. Wherein, the characteristic of the first cache queue is that the first pipeline can start the next execution after pushing the data.

[0054] S103, respectively on a plurality of second pipelines, each first pipeline output of the processing result is read from the first cache queue for the second parallel processing, to complete the inference of the plurality of offline models.

[0055] In the embodiment, the second pipeline refers to a pipeline for processing the processing result output by the first pipeline. Specifically, one processing result output by the first pipeline can be processed on each second pipeline, or multiple processing results output by the first pipeline can be processed on each second pipeline. The processing process includes pre-processing, inference calculation, post-processing, and the like. The types of the multiple second pipelines can be the same or different.

[0056] In the embodiment, since the first cache queue is characterized in that the first pipeline can start the next execution after the first pipeline pushes data, when the first cache queue stores the processing result output by the first pipeline, the processor can simultaneously read the processing result output by the first pipeline from the first cache queue on the multiple second pipelines respectively, and perform parallel processing on each second pipeline on the processing result read by each second pipeline. Specifically, the processor can simultaneously read all the processing results output by the first pipeline from the first cache queue on the multiple second pipelines respectively, to perform parallel processing on each second pipeline on all the processing results read by each second pipeline. Alternatively, the processor can read one processing result output by the first pipeline from the first cache queue on the multiple second pipelines in sequence, and perform parallel processing on each second pipeline on the processing result read by each second pipeline after data is obtained on all the second pipelines. Alternatively, the processor can read all the processing results output by the first pipeline from the first cache queue on the multiple second pipelines in sequence, and perform parallel processing on each second pipeline on all the processing results read by each second pipeline after data is obtained on all the second pipelines. When the data processing on each second pipeline is completed, the inference process on the to-be-processed data is completed, and the processor can obtain new to-be-processed data, and then perform inference on the new to-be-processed data on the first pipeline, the first cache queue, and the second pipeline according to the method of S101-S103, until there is no new to-be-processed data to be inferred, i.e., there is no to-be-processed data to be processed on the first pipeline, there is no data in the first cache queue, and the data on the second pipeline is processed, so that the inference of the multiple models is completed.

[0057] In the method for offline model inference, multiple first pipelines and multiple second pipelines are arranged to parallelize the processing data, and the second pipeline and the first pipeline are arranged in cascade, i.e., the second pipeline processes the processing result output by the first pipeline, so that the data is processed by different pipelines in a pipelining manner. Based on this, the method for offline model inference provided by the application realizes a parallelization and pipelining operation of multiple pipelines in the application scenario of offline model inference, and greatly improves the speed and efficiency of model inference compared with the method for performing inference calculation by a single pipeline.

[0058] In actual applications, the plurality of offline models loaded by the processor when performing offline model inference can be offline models of the same type, such as a plurality of offline models being offline models of convolution operation; the plurality of offline models loaded can also be offline models of different types, such as one offline model being an offline model of convolution operation and another offline model being an offline model of pooling operation. When the plurality of offline models loaded by the processor when performing offline model inference are offline models of the same type, the processor, when performing the step of S102, specifically performs: performing first parallel processing on the to-be-processed data of the plurality of offline models on the plurality of first pipelines respectively, and storing the processing results output by each first pipeline in the first cache queue in sequence according to the order of outputting the processing results by each first pipeline.

[0059] The processing result output by the first pipeline is a processing result obtained by the processor after performing preprocessing, inference calculation, post-processing, and the like on the to-be-processed data on the first pipeline. When performing first parallel processing on the to-be-processed data of the plurality of offline models on the plurality of first pipelines, the processing speed can be the same or can be different. Therefore, when the processing speed of performing first parallel processing on the to-be-processed data of the plurality of models on the plurality of first pipelines is different, the processor can store the processing results output by each first pipeline in the first cache queue in sequence according to the order of outputting the processing results by each first pipeline; when the processing speed of performing first parallel processing on the to-be-processed data of the plurality of offline models on the plurality of first pipelines is the same, the processor can store the processing results output by each first pipeline in the first cache queue in any order.

[0060] When there is a processing result output by a first pipeline in the first storage queue, the processor can perform the step of S103, and the specific execution process is: reading at least one processing result output by a first pipeline from the first cache queue on each second pipeline respectively to perform second parallel processing.

[0061] In this embodiment, the number of first pipelines can be consistent with the number of second pipelines, or the number of first pipelines can be inconsistent with the number of second pipelines. When the number of first pipelines is consistent with the number of second pipelines, and the first cache queue stores processing results output by multiple first pipelines, the processor can read the processing results output by one first pipeline in the first cache queue on each second pipeline, as long as different processing results output by first pipelines are read on each second pipeline, and then the processing results read on each second pipeline are processed in parallel. For example, assuming that there are three second pipelines: #1 second pipeline, #2 second pipeline, and #3 second pipeline, and the first cache queue stores processing results output by #1 first pipeline, processing results output by #2 first pipeline, and processing results output by #3 first pipeline, the #1 second pipeline can read the processing results output by #1 first pipeline from the first cache queue, the #2 second pipeline can read the processing results output by #2 first pipeline from the first cache queue, and the #3 third pipeline can read the processing results output by #3 first pipeline from the first cache queue. Optionally, the #1 second pipeline can read the processing results output by #2 first pipeline from the first cache queue, the #2 second pipeline can read the processing results output by #3 first pipeline from the first cache queue, and the #3 second pipeline can read the processing results output by #1 first pipeline from the first cache queue. Optionally, the #1 second pipeline can read the processing results output by #3 first pipeline from the first cache queue, the #2 second pipeline can read the processing results output by #1 first pipeline from the first cache queue, and the #3 third pipeline can read the processing results output by #2 first pipeline from the first cache queue.

[0062] Optionally, when the number of first pipelines is not the same as the number of second pipelines, and the first buffer queue stores multiple processing results output from the first pipelines, the processor can read multiple processing results output from the first buffer queue on each second pipeline. The number of processing results read by each pipeline can be the same or different, and then the processor processes the processing results read by each pipeline in parallel. For example, assuming there are two second pipelines: #1 and #2, and the first buffer queue stores the processing results output from #1, #2, and #3, then on the #1 second pipeline, the processing results output from #1 and #2 can be read from the first buffer queue; on the #2 second pipeline, the processing result output from #3 can be read from the first buffer queue. For example, suppose there are two second pipelines: #1 and #2. The first buffer queue stores the processing results of the first pipelines #1, #2, #3, and #4. Then, on the second pipeline #1, the processing results of the first pipelines #1 and #2 can be read from the first buffer queue; and on the second pipeline #2, the processing results of the first pipelines #3 and #4 can be read from the first buffer queue.

[0063] This application exemplarily illustrates the reasoning method for the same type of model described in the above embodiments when the number of the first pipeline is the same as the number of the second pipeline. Figure 3 The model inference structure shown assumes three first pipelines: Pipeline A1, Pipeline A2, and Pipeline A3. These three pipelines are of the same type and run in parallel, with their results stored in the Basic Queue without needing to be synchronized. Similarly, there are three second pipelines: Pipeline B1, Pipeline B2, and Pipeline B3. These second pipelines are also of the same type and run in parallel, each independently reading data from the Basic Queue. The specific calculation process is as follows:

[0064] The processor first creates a first buffer queue, the Basic Queue, which is positioned between the parallel pipelines A1, A2, and A3 and the parallel pipelines B1, B2, and B3. The processing results from the parallel computations of pipelines A1, A2, and A3 are placed into the Basic Queue. The Basic Queue is characterized by the following behavior: pipelines A1, A2, and A3 can begin the next execution after pushing data; pipelines B1, B2, and B3 sequentially push a copy of the Basic Queue. The data in the Queue includes one of the following: the processing result output from Pipeline A1, Pipeline A2, and Pipeline A3. For example, Pipeline B1 pops the processing result output from Pipeline A1 from the Basic Queue, Pipeline B2 pops the processing result output from Pipeline B2 from the Basic Queue, and Pipeline B3 pops the processing result output from Pipeline A3 from the Basic Queue.

[0065] The processor runs start, at this time the first cache queue Basic Queue is empty, therefore, the second pipeline Pipeline B1, the second pipeline Pipeline B2, the second pipeline Pipeline B3 are in a waiting state. When the model data input, the first pipeline Pipeline A1, the first pipeline Pipeline A2, the first pipeline Pipeline A3 trigger execution. After starting parallel processing data on the first pipeline Pipeline A1, the first pipeline Pipeline A2, the first pipeline Pipeline A3, and the processing results are stored in the first cache queue Basic Queue, the second pipeline Pipeline B1, the second pipeline Pipeline B2, the second pipeline Pipeline B3 are triggered, after the data is popped from the first cache queue Basic Queue on the second pipeline Pipeline B1, the second pipeline Pipeline B2, the second pipeline Pipeline B3, the first pipeline Pipeline A1, the first pipeline Pipeline A2, the first pipeline Pipeline A3 and the second pipeline Pipeline B1, the second pipeline Pipeline B2, the second pipeline Pipeline B3 execute concurrently.

[0066] After the first pipeline Pipeline A1, the first pipeline Pipeline A2, the first pipeline Pipeline A3 process all data, at this time, the first pipeline Pipeline A1, the first pipeline Pipeline A2, the first pipeline Pipeline A3 end execution, then, when there is no data in the first cache queue Basic Queue, the second pipeline Pipeline B1, the second pipeline Pipeline B2, the second pipeline Pipeline B3 also end running, the inference process ends.

[0067] Optionally, the exemplary description of the present application is in the case that the number of the first pipeline and the number of the second pipeline are inconsistent, the inference method of the same type model described in the above embodiment, such as Figure 4The shown model inference structure assumes that there are three first pipelines, i.e., a first pipeline Pipeline A4, a first pipeline Pipeline A5, and a first pipeline Pipeline A6, and the first pipeline Pipeline A4, the first pipeline Pipeline A5, and the first pipeline Pipeline A6 are pipelines of the same type, run in parallel, and the running results of each pipeline do not need to be synchronized and can be stored into a first cache queue Basic Queue. The shown model inference structure assumes that there are two second pipelines, i.e., a second pipeline Pipeline B4 and a second pipeline Pipeline B5, and the second pipeline Pipeline B4 and the second pipeline Pipeline B5 are also pipelines of the same type, run in parallel, and independently read data from the first cache queue Basic Queue. The specific calculation process is as follows:

[0068] The processor creates a first cache queue Basic Queue, which is set between the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 arranged in parallel and the second pipeline Pipeline B4 and the second pipeline Pipeline B5 arranged in parallel. The processing results output by the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 after parallel calculation are put into the first cache queue Basic Queue. The first cache queue Basic Queue is characterized in that it can start the next execution after the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 push data. The second pipeline Pipeline B4 and the second pipeline Pipeline B5 Pop the processing results output by the first pipeline in turn from the first cache queue Basic Queue. Specifically, the second pipeline Pipeline B4 can Pop the processing results output by the first pipeline Pipeline A4 and the first pipeline Pipeline A5 from the first cache queue Basic Queue, and the second pipeline Pipeline B5 can Pop the processing results output by the first pipeline Pipeline A6 from the first cache queue Basic Queue. Optionally, the second pipeline Pipeline B4 can also Pop the processing results output by the first pipeline Pipeline A4 from the first cache queue Basic Queue, and the second pipeline Pipeline B5 can Pop the processing results output by the first pipeline Pipeline A5 and the first pipeline Pipeline A6 from the first cache queue Basic Queue. It should be noted that the second pipeline Pipeline B4 and the second pipeline Pipeline B5 can each Pop any combination of the processing results output by the first pipeline from the first cache queue Basic Queue, which is not limited here.

[0069] The processor runs start, at this time the first cache queue Basic Queue is empty, therefore, the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 are in the waiting state. When the model data input, the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 trigger execution. After starting parallel processing data on the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6, and the processing result is stored in the first cache queue Basic Queue, the second pipeline Pipeline B4 and the second pipeline Pipeline B5 are triggered, after the data is popped from the first cache queue Basic Queue on the second pipeline Pipeline B4 and the second pipeline Pipeline B5, the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 and the second pipeline Pipeline B4 and the second pipeline Pipeline B5, concurrent execution.

[0070] After the first pipeline Pipeline A4, the first pipeline Pipeline A5, the first pipeline Pipeline A6 process all data, at this time, the second pipeline Pipeline A4, the second pipeline Pipeline A5, the second pipeline Pipeline A6 end execution, then, when there is no data in the first cache queue Basic Queue, the second pipeline Pipeline B4 and the second pipeline Pipeline B5 also end running, the inference process ends.

[0071] Need to explain, Figure 3 And Figure 4 The model inference structure in the embodiment, wherein when the types of the respective first pipelines are the same, the types of the respective second pipelines can be the same or different, Figure 3 And Figure 4 It is only an example of a second pipeline of the same type, which is not limited here. When the types of the second pipelines are different, the model inference structure described in the embodiment can also be used Figure 3 Or Figure 4 The model inference structure described in the embodiment is used for data inference, and the specific calculation process is consistent with the calculation process described above, which is not repeated here.

[0072] When the processor performs offline model inference, the plurality of offline models loaded are offline models of different types, the first cache queue includes a first storage queue and a first copy queue, and the processor performs the step of S102, likeFigure 5 As shown, the "storing the processing results of each first pipeline to the first cache queue" in S102 specifically includes the following steps:

[0073] S201, synchronizing the processing results output by each first pipeline, and storing the synchronized processing results output by each first pipeline to the first storage queue.

[0074] In this embodiment, after the first parallel processing of the to-be-processed data of the plurality of offline models on the plurality of first pipelines, the processing results output by each first pipeline need to be synchronized, so that the processor simultaneously obtains the processing results output by the first pipeline, and then the synchronized processing results output by all first pipelines are stored to the first storage queue.

[0075] S202, copying the processing results output by each first pipeline in the first storage queue according to a preset number of copies to obtain a preset number of copy results, and storing the preset number of copy results to the first copy queue.

[0076] The preset number of copies is equal to the number of second pipelines, and the preset number of copies can be set by the processor in advance according to the number of second pipelines. For example, if there are three second pipelines, the corresponding preset number of copies is three, that is, the processing results output by each first pipeline in the first storage queue are copied three times, and each copy result will include the processing results of each first pipeline.

[0077] In this embodiment, after the processor stores the synchronized processing results output by each first pipeline to the first storage queue, the processing results in the first storage queue can be further copied according to the preset number of copies to obtain a preset number of copy results, so that each copy result includes the processing results of each second pipeline, and then the preset number of copy results are stored to the first copy queue according to the arrangement mode of the number of copies.

[0078] After the processor stores the copy results to the first copy queue, the processor can execute the steps of S103, specifically can execute: reading a copy result from the first copy queue on each second pipeline for second parallel processing.

[0079] In this embodiment, the number of first pipelines can be consistent with the number of second pipelines, and the number of first pipelines can also be inconsistent with the number of second pipelines. The first copy queue stores multiple copies of results, that is, multiple processing results output by the first pipelines. The processor can read one copy of results from the first copy queue on each second pipeline for processing. For example, in the case where the number of first pipelines is the same as the number of second pipelines, assuming that there are three first pipelines, #1 first pipeline, #2 first pipeline, and #3 first pipeline; there are three second pipelines, #1 second pipeline, #2 second pipeline, and #3 second pipeline; the first copy queue stores three copies of results, each of which includes the processing result output by the #1 first pipeline, the processing result output by the #2 second pipeline, and the processing result output by the #3 third pipeline; then one copy of results is read from the first copy queue on the #1 second pipeline, one copy of results is read from the first copy queue on the #2 second pipeline, and one copy of results is read from the first copy queue on the #3 second pipeline. The copies of results read from the first copy queue on the #1 second pipeline, the #2 second pipeline, and the #3 second pipeline are the same.

[0080] For another example, in the case where the number of first pipelines is inconsistent with the number of second pipelines, assuming that there are three first pipelines, #1 first pipeline, #2 first pipeline, and #3 first pipeline; there are two second pipelines, #1 second pipeline and #2 second pipeline; the first copy queue stores two copies of results, each of which includes the processing result output by the #1 first pipeline, the processing result output by the #2 second pipeline, and the processing result output by the #3 third pipeline; then one copy of results is read from the first copy queue on the #1 second pipeline, and one copy of results is read from the first copy queue on the #2 second pipeline. The copies of results read from the first copy queue on the #1 second pipeline and the #2 second pipeline are the same.

[0081] The exemplary description of the present application is in the case where the number of first pipelines is consistent with the number of second pipelines. The inference method of different types of models described in the above embodiments, such as Figure 6As shown, assuming there are three first pipelines, i.e., the first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1, and the first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1 are pipelines of different types, they run in parallel, and the running results of each of them need to be synchronized before being put into the first storage queue (Sync Queue) in the first cache queue. Assuming there are three second pipelines, i.e., the second pipeline Pipeline D1, the second pipeline Pipeline E1, and the second pipeline Pipeline F1, and the second pipeline Pipeline D1, the second pipeline Pipeline E1, and the second pipeline Pipeline F1 are also pipelines of different types, they run in parallel, and each of them takes the same data from the first copy queue (Copy Queue) for calculation, wherein the data includes the processing result output by the first pipeline Pipeline A1, the processing result output by the first pipeline Pipeline B1, and the processing result output by the first pipeline Pipeline C1. The specific calculation process is as follows:

[0082] A first cache queue is created and arranged between first pipeline A1, first pipeline B1, first pipeline C1 in parallel arrangement and second pipeline D1, second pipeline E1, second pipeline F1 in parallel arrangement. The results of the parallel calculation of first pipeline A1, first pipeline B1, first pipeline C1 are put into the first cache queue. The first cache queue is characterized in that it is composed of two queues inside, one is a first storage queue Sync Queue, and the other is a first copy queue Copy Queue. The input port of the first copy queue Copy Queue is connected to the output port of the first storage queue Sync Queue. In actual application, the synchronization between the first storage queue Sync Queue and the first copy queue Copy Queue can be handled by an additional thread. After all the data in the first storage queue Sync Queue is popped, the thread can notify first pipeline A1, first pipeline B1, first pipeline C1 to start the next execution, and copy all the popped data into the first copy queue Copy Queue three times. Second pipeline D1, second pipeline E1, second pipeline F1 can pop data from the first copy queue Copy Queue in turn. When the last second pipeline pops data from the first copy queue Copy Queue, it notifies the thread managing the first storage queue Sync Queue and the first copy queue Copy Queue to perform the next data synchronization.

[0083] The running starts, at this time the first cache queue is empty, therefore, the second pipeline D1, the second pipeline E1, the second pipeline F1 are in the waiting state. The first pipeline A1, the first pipeline B1, the first pipeline C1 trigger execution, after the first pipeline A1, the first pipeline B1, the first pipeline C1 have data put into the first cache queue, the second pipeline D1, the second pipeline E1, the second pipeline F1 are triggered, the first pipeline A1, the first pipeline B1, the first pipeline C1 and the second pipeline D1, the second pipeline E1, the second pipeline F1 execute concurrently.

[0084] After the first pipeline A1, the first pipeline B1, the first pipeline C1 process all data, the first pipeline A1, the first pipeline B1, the first pipeline C1 end running, when there is no data in the first cache queue, the second pipeline D1, the second pipeline E1, the second pipeline F1 also end running, the inference process ends.

[0085] Optionally, the exemplary description of the present application is in the case that the number of the first pipeline is not consistent with the number of the second pipeline, the inference method of different types of models described in the above embodiment, such as Figure 7The shown model inference structure assumes that there are three first pipelines, i.e., a first pipeline Pipeline A1, a first pipeline Pipeline B1, and a first pipeline Pipeline C1, and the first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1 are pipelines of different types, which run in parallel and the running results of each of the pipelines need to be synchronized before being put into a first storage queue (Sync Queue) in a first cache queue.

[0086] A first cache queue is created and is arranged between first pipelines Pipeline A1, Pipeline B1, Pipeline C1 arranged in parallel and second pipelines Pipeline D1 and Pipeline E1 arranged in parallel. The results of the parallel computation of the first pipelines Pipeline A1, Pipeline B1, Pipeline C1 are placed in the first cache queue. The first cache queue is characterized in that it is composed of two queues, a first storage queue Sync Queue and a first copy queue Copy Queue. The input of the first copy queue Copy Queue is connected to the output of the first storage queue Sync Queue. In actual application, the synchronization between the first storage queue Sync Queue and the first copy queue Copy Queue can be handled by an extra thread. After all data in the first storage queue Sync Queue is popped, the thread can notify the first pipelines Pipeline A1, Pipeline B1, Pipeline C1 to start the next execution. All the popped data is copied twice and placed in the first copy queue Copy Queue. The second pipelines Pipeline D1 and Pipeline E1 can pop data from the first copy queue Copy Queue in turn. When the last second pipeline pops data from the first copy queue Copy Queue, the thread managing the first storage queue Sync Queue and the first copy queue Copy Queue is notified to perform the next data synchronization.

[0087] When the running starts, the first cache queue is empty, so the second pipelines Pipeline D1 and Pipeline E1 are in a waiting state. The first pipelines Pipeline A1, Pipeline B1, Pipeline C1 are triggered to execute. When the first pipelines Pipeline A1, Pipeline B1, Pipeline C1 have data placed in the first cache queue, the second pipelines Pipeline D1 and Pipeline E1 are triggered. The first pipelines Pipeline A1, Pipeline B1, Pipeline C1 and the second pipelines Pipeline D1 and Pipeline E1 execute in parallel.

[0088] When the first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1 finish processing all data, the first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1 end running, and when there is no data in the first cache queue, the second pipeline Pipeline D1 and the second pipeline Pipeline E1 also end running, and the inference process ends.

[0089] Need to be explained, Figure 6 And Figure 7 The model inference structure in the embodiment, wherein when the types of the first pipelines are the same, the types of the corresponding second pipelines can be the same or different, Figure 6 And Figure 7 It is only an example of a second pipeline of the same type, which is not limited here. When the types of the second pipelines are different, the model inference structure described in the embodiment can also be used for data inference, and the specific calculation process is consistent with the calculation process described above, which is not described here. Figure 6 Or Figure 7 The model inference structure described in the embodiment, and the specific calculation process is consistent with the calculation process described above, which is not described here.

[0090] In the above embodiment, the pipelines corresponding to the same or different types of models can be run in parallel through multi-threading, and the pipelines are run in parallel through the Queue, and the length of the Queue depends on the number of pipelines of different types, that is, the more the number of the first pipelines in the embodiment, the longer the length of the Queue. The above embodiment simultaneously realizes the parallelization and pipelining of the model inference, and when facing multiple offline models in the application scenarios of machine learning network operation or neural network operation, the inference speed and efficiency can be greatly improved.

[0091] The application also provides a parallelization and pipelining method for improving the performance of a single pipeline. The following embodiments will be described for a single pipeline. Each pipeline includes a plurality of first threads, a second thread, and a plurality of third threads, wherein the pipeline is a first pipeline and / or a second pipeline. Figure 8 A flowchart of processing of data to be processed by a single pipeline of an embodiment of the application on an offline model is shown.

[0092] S301, a plurality of first threads are called to perform parallel preprocessing on data to be processed by a pipeline corresponding to an offline model, and preprocessing results of each first thread are stored in a second cache queue.

[0093] The pre-processing indicates a series of pre-processing work required before data enters a model for inference, such as normalizing, filtering, denoising, dimension conversion, and the like. The first thread is used to load data and pre-process the loaded data, and therefore the type of the first thread in this embodiment refers to the type of logic for processing data on the first thread, such as a first first thread for filtering data, and a second first thread for denoising data, and therefore the types of the two first threads are different. In actual application, the types of the plurality of first threads can be the same or different.

[0094] In this embodiment, taking the first pipeline as an example, when the processor performs first parallel processing on the to-be-processed data of an offline model on the first pipeline, the processor can set a plurality of threads inside the first pipeline to perform parallel loading on the to-be-processed data of a model according to actual application requirements. Specifically, the processor can divide the to-be-processed data of an offline model into a plurality of thread task data according to the number of threads, and then call each first thread to perform parallel pre-processing on each thread task data, thereby completing the parallel pre-processing of the to-be-processed data of an offline model by the plurality of first threads. It should be noted that the size of the thread task data processed by each first thread can be the same or different. For example, assuming that there are three first threads, and the to-be-processed data is 120 image data of the same size, the to-be-processed data can be divided into three thread task data, that is, the 120 images can be divided into three thread task data, and each thread task data can include 40 images, or each thread task data can include different numbers of images (such as 30, 40, or 50). After the processor calls the plurality of first threads to perform parallel pre-processing on the to-be-processed data of an offline model, the pre-processing results of each first thread can be obtained, and then a second cache queue can be created, and the pre-processing results of each first thread can be stored in the second cache queue. Alternatively, the processor can create the second cache queue before performing the pre-processing of the first thread, and then store the pre-processing results of each first thread in the created second cache queue after performing the pre-processing of the first thread. When the processor obtains the pre-processing results of each first thread in sequence, the processor can store the pre-processing results of each first thread in the second cache queue in sequence according to the order in which the data is pre-processed by the first thread; alternatively, when the processor obtains the pre-processing results of each first thread at the same time, the processor can store the pre-processing results of each first thread in the second cache queue together.

[0095] S302, calling the second thread to load the offline model, and performing inference on the offline model based on each pre-processing result in the second cache queue, obtaining the model inference result of each second thread, and storing the model inference result of each second thread in the third cache queue.

[0096] In the embodiment, after the pre-processing results of the first threads are stored in the second cache queue, the processor can call the second thread to load the offline model first, and then read the pre-processing results of the first threads from the second cache queue as input data of the loaded offline model, and then input the read pre-processing results of the first threads into the loaded offline model for inference calculation to obtain the model inference results in the second thread. In actual application, the processor calls the second thread to input the pre-processing results of the first threads into the loaded offline model respectively for inference calculation to obtain the model inference results in the second thread. Alternatively, the processor can call the second thread to input the pre-processing results of the first threads into the loaded offline model together for inference calculation to obtain the model inference results in the second thread respectively. After the processor calls the second thread to perform offline model inference based on the pre-processing results in the second cache queue to obtain the model inference results in the second thread, a third cache queue can be created, and the model inference results in the second thread obtained by model inference can be stored in the third cache queue. Alternatively, the processor can create the third cache queue before performing the model inference of the second thread, and then store the model inference results of the second thread in the third cache queue after performing the model inference of the second thread. Specifically, when the processor obtains the model inference results of the second thread, the processor can store the model inference results of the second thread in the third cache queue in the order of output of the model inference results of the second thread. If the processor obtains the model inference results of the second thread at the same time, the processor can store the model inference results of the second thread in the third cache queue together regardless of the order.

[0097] S303, calling a plurality of third threads to read the model inference results of the second threads from the third cache queue for parallel post-processing.

[0098] The post-processing represents a series of post-processing work required by the data obtained after model inference, for example, the post-processing work can be a series of frame marking, text marking, drawing and other complex scalar type calculations. For example, when the offline model inference is used for image classification or recognition, the size of the target object in the image needs to be marked by frame or text, and the corresponding post-processing is used to mark the target object by frame or text. When the offline model inference is used for image segmentation, image reconstruction needs to be performed based on the segmented image data, and the corresponding post-processing is used to perform drawing operation. The third thread is used for post-processing of the data. The type of the third thread refers to the type of logic for processing the data. For example, when the offline model inference is used for image classification or recognition, the third thread is called to mark the classified or recognized target object by frame or text, so as to mark the classified or recognized target object. When the offline model inference is used for image segmentation, the third thread is called to draw based on the segmented image data, so as to obtain the segmented image. In actual application, the types of the plurality of third threads can be the same or different according to actual processing requirements.

[0099] In this embodiment, after the processor calls the second thread to complete the model inference calculation and stores the model inference results of each second thread in the third cache queue, the processor can call multiple third threads to read the model inference results of one second thread from the third cache queue respectively, and each third thread performs parallel post-processing on the model inference result read by itself. Alternatively, the processor can also call multiple third threads to read all model inference results of all second threads from the third cache queue at the same time, and each third thread performs parallel post-processing on all model inference results read by itself. Alternatively, the processor can also call multiple third threads to read the model inference results of one second thread from the third cache queue in sequence, and after all third threads obtain the data, each third thread performs parallel post-processing on the model inference result read by itself. Alternatively, the processor can also call multiple third threads to read all model inference results of all second threads from the third cache queue in sequence, and after all third threads obtain the data, each third thread performs parallel post-processing on all model inference results read by itself. When each third thread completes the processing of the data read by itself, the processing process of the first pipeline on the to-be-processed data is completed, and then the processor can obtain new to-be-processed data, and call multiple first threads, one second thread, a second cache queue, a third cache queue, and multiple third threads to process the new to-be-processed data according to the method of S301-S303, until there is no new to-be-processed data to be processed, that is, there is no data to be loaded and pre-processed in the first thread, there is no data in the second cache queue, there is no data in the third cache queue, there is no data to be inferred in the second thread, and there is no data to be post-processed in the third thread. It can be considered that the processing work of one first pipeline is completed, and when each first pipeline is executed according to the method of S301-S303, the first parallel processing of multiple first pipelines is completed.

[0100] In actual applications, multiple types of processors are installed on the processor, which are used to perform model inference calculation according to the method described in the foregoing embodiments. For example, at least one of a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a tensor processing unit (TPU), etc. can be installed on the processor at the same time. Based on this, the multiple first threads, the second thread, and the multiple third threads can run on the same processor or different processors.

[0101] In an embodiment, the first plurality of threads and the third plurality of threads are executed on a first processor, and the second thread is executed on a second processor. The first processor can be a CPU, and the second processor can be an NPU. When the first plurality of threads and the third plurality of threads are executed on the CPU, it means that the CPU is responsible for the pre-processing and post-processing of data. When the second thread is executed on the NPU, it means that the NPU is responsible for the inference of the network model. It is a great challenge to map all the complex calculations in the deep learning network model to the NPU chip. On the one hand, the NPU chip is not suitable for a large number of complex scalar type calculations, which is reflected in poor performance. On the other hand, the types of operators in the network are increasingly rich, and the adaptation of these operators on the NPU is more costly than on the CPU. However, the increasingly rich deployment scenarios of deep learning networks urgently need fast forward inference. Therefore, the work belonging to complex scalar type calculation (pre-processing and post-processing of data) is placed on the CPU for operation, and the work belonging to network model inference calculation (loading model for calculation) is placed on the NPU for operation, which can fully utilize the advantages of CPU and NPU to speed up the inference efficiency and speed of offline models with optimal resources. Moreover, the first plurality of threads, the second cache queue, the third cache queue, and the third plurality of threads form a parallel and pipelined way to process data, further improving the model inference speed and efficiency.

[0102] In actual applications, the processor calls the first plurality of threads to execute data loading and pre-processing. The logic for processing data on the first plurality of threads can be the same or different. When the logic for processing data on the first plurality of threads is the same, the processor, when executing the step of S302, specifically executes: storing the pre-processing results of each first thread to the second cache queue in sequence according to the order in which each first thread outputs the pre-processing results.

[0103] The speeds of the first plurality of threads for pre-processing the plurality of loaded data can be the same or different. When the speeds of the first plurality of threads for pre-processing the plurality of data are different, the processor can store the pre-processing results of each first thread to the second cache queue in sequence according to the order in which each first thread outputs the pre-processing results. When the speeds of the first plurality of threads for pre-processing the plurality of data are the same, the processor can store the pre-processing results of each first thread to the second cache queue in any order.

[0104] When the processor stores the preprocessing results of each first thread into the second cache queue, the processor can execute the steps of S302, specifically can execute: calling the second thread to load the offline model, and sequentially reading the preprocessing results of each first thread from the second cache queue, and then inputting the read preprocessing results into the loaded offline model in sequence for inference to obtain the model inference results of each second thread; optionally, the processor can also input all the preprocessing results of the first threads into the loaded offline model from the second cache queue for inference to obtain the model inference results of each second thread.

[0105] When the processor stores the model inference results of each second thread into the third cache queue, the processor can execute the steps of S303, specifically can execute: calling each third thread to read at least one model inference result of a second thread from the third cache queue for parallel post-processing.

[0106] In this embodiment, the number of first threads can be consistent with the number of third threads, and the number of first threads can also be inconsistent with the number of third threads. When the number of first threads is consistent with the number of third threads, and the third cache queue stores a plurality of model inference results of second threads, the processor can call each third thread to read the model inference result of a second thread in the third cache queue respectively, as long as each third thread reads the model inference result of a different second thread, and then parallel post-processes the model inference result read by each third thread. For example, assuming that there are three first threads: #1 first thread, #2 first thread and #3 first thread, three third threads: #1 third thread, #2 third thread and #3 third thread, and the third cache queue stores the model inference result of #1 second thread obtained by the preprocessing result of #1 first thread on the second thread after model inference, the model inference result of #2 second thread obtained by the preprocessing result of #2 first thread on the second thread after model inference, and the model inference result of #3 second thread obtained by the preprocessing result of #3 first thread on the second thread after model inference, #1 third thread can read the model inference result of #1 second thread from the third cache queue; #2 third thread can read the model inference result of #2 second thread from the third cache queue; and #3 third thread can read the model inference result of #3 second thread from the third cache queue. Optionally, #1 third thread can read the model inference result of #2 second thread from the third cache queue; #2 third thread can read the model inference result of #3 second thread from the third cache queue; and #3 third thread can read the model inference result of #1 second thread from the third cache queue. Optionally, #1 third thread can read the model inference result of #3 second thread from the third cache queue; #2 third thread can read the model inference result of #1 second thread from the third cache queue; and #3 third thread can read the model inference result of #2 second thread from the third cache queue.

[0107] Optionally, when the number of first threads is inconsistent with the number of third threads, and the third cache queue stores a plurality of model inference results of second threads, the processor can call each third thread to read the plurality of model inference results of second threads in the third cache queue respectively, and then parallelly post-process the model inference results read respectively. For example, assuming that there are three first threads: #1 first thread, #2 first thread and #3 first thread, two third threads: #1 third thread and #2 third thread, and the third cache queue stores the model inference result of #1 second thread obtained by the pre-processing result of #1 first thread on the second thread after model inference, the model inference result of #2 second thread obtained by the pre-processing result of #2 first thread on the second thread after model inference, and the model inference result of #3 second thread obtained by the pre-processing result of #3 first thread on the second thread after model inference, #1 third thread can read the model inference result of #1 second thread and the model inference result of #2 second thread from the third cache queue, and #2 third thread can read the model inference result of #3 second thread from the third cache queue. It should be noted that each third thread can read out any combination of model inference results of second threads from the third cache queue, which is not limited here.

[0108] The exemplary description of the present application is that when the number of first threads is consistent with the number of third threads, the data processing method on the pipeline corresponding to the same logic of processing data on each first thread described in the above embodiment is as follows: Figure 9As shown, there are three first threads, i.e., first thread Data Loader A1, first thread Data Loader A2, and first thread Data Loader A3, and the types of the first thread Data Loader A1, first thread Data Loader A2, and first thread Data Loader A3 are the same. There are three third threads, i.e., third thread Post Processor A1, third thread Post Processor A2, and third thread Post Processor A3, and the types of the third thread Post Processor A1, third thread Post Processor A2, and third thread Post Processor A3 are the same. The first thread Data Loader A1, first thread Data Loader A2, and first thread Data Loader A3 are executed in multi-thread concurrency. The third thread Post Processor A1, third thread Post Processor A2, and third thread Post Processor A3 are executed in multi-thread concurrency. A model operation unit (Model Runner) is created between the first thread Data Loader and the third thread Post Processor, a second cache queue Basic Queue 1 is created between the first thread Data Loader and the model operation unit Model Runner, and a third cache queue Basic Queue 2 is created between the model operation unit Model Runner and the third thread Post Processor. As shown in FIG. 9, the first thread Data Loader A1, first thread Data Loader A2, and first thread Data Loader A3 are the same type of Data Loader, they run in parallel, and the running results of each of them can be put into the second cache queue Basic Queue 1 without synchronization. The third thread Post Processor A1, third thread Post Processor A2, and third thread Post Processor A3 are the same type of Post Processor, they run in parallel, and each of them independently takes data from the third cache queue Basic Queue 2. The specific calculation process is as follows:

[0109] A second cache queue Basic Queue1 is created and arranged between the first thread Data Loader A1, the first thread Data Loader A2, the first thread Data Loader A3 and the model operation unit Model Runner in parallel arrangement, the results of the parallel calculation of the first thread Data Loader A1, the first thread Data Loader A2, the first thread Data Loader A3 are put into the second cache queue Basic Queue1, the second cache queue Basic Queue1 is characterized by starting the next execution as soon as a first thread Data Loader pushes data, and the model operation unit Model Runner pops one piece of data in the second cache queue Basic Queue1 at a time. A third cache queue Basic Queue2 is created and arranged between the model operation unit Model Runner and the third thread Post Processor A1, the third thread Post Processor A2, the third thread Post Processor A3 in parallel arrangement, the third cache queue Basic Queue2 is characterized by the model operation unit Model Runner pushing one piece of data to the third cache queue Basic Queue2 at a time, and the third thread Post Processor popping one piece of data from the third cache queue Basic Queue2 at a time. Wherein, one piece of data includes the model inference result obtained by the pre-processing result of the first thread DataLoader A1 after model inference on the model operation unit Model Runner, or the model inference result obtained by the pre-processing result of the first thread Data Loader A2 after model inference on the model operation unit Model Runner, or the model inference result obtained by the pre-processing result of the first thread Data Loader A3 after model inference on the model operation unit Model Runner.

[0110] The running starts, at this time, the second cache queue Basic Queue 1 and the third cache queue Basic Queue 2 are empty, therefore, the model operation unit Model Runner is in a waiting state with the third thread Post Processor A1, the third thread Post Processor A2 and the third thread Post Processor A3. The first thread Data Loader A1, the first thread Data Loader A2 and the first thread Data Loader A3 trigger execution. After the first thread Data Loader has data put into the second cache queue Basic Queue 1, the model operation unit Model Runner is triggered, after the model operation unit Model Runner pops the data, the first thread Data Loader and the model operation unit Model Runner execute concurrently. After the model operation unit Model Runner runs, the model operation unit Model Runner pushes the data to the third cache queue Basic Queue 2 associated with the third thread Post Processor, the third thread Post Processor triggers running, and the model operation unit Model Runner and the third thread Post Processor execute concurrently.

[0111] After the first thread Data Loader processes all data, there is no data in the second cache queue Basic Queue 1 and the third cache queue Basic Queue 2, at this time, the model operation unit Model Runner ends execution, then the third thread Post Processor also ends running. The inference process ends.

[0112] It should be noted that when the first thread Data Loader outputs data in the above embodiment, the processor can call the second thread to store the data output by the first thread into the second cache queue; when the model operation unit Model Runner in the above embodiment executes, the processor can also call the second thread to implement; when the model operation unit Model Runner in the above embodiment runs to output the model inference results of each second thread, the processor can also call the second thread to store the model inference results of each second thread into the third cache queue. The above multiple Data Loaders correspond to multiple first threads, and the multiple Post Processors correspond to multiple third threads, that is, the Data Loader and the Post Processor run on the first processor, and the Model Runner runs on the second processor. In actual application, the first processor is a CPU, and the second processor is an NPU.

[0113] Optionally, the data processing method of the corresponding pipeline when the number of the first pipelines and the number of the second pipelines are inconsistent is the same as the logical processing of the data on each first thread as described in the above embodiments, such as Figure 10 As shown in the figure, it is assumed that there are three first threads, i.e., the first thread Data Loader A1, the first thread Data Loader A2, and the first thread Data Loader A3, and the types of the first thread Data Loader A1, the first thread Data Loader A2, and the first thread Data Loader A3 are the same. There are two third threads, i.e., the third thread Post Processor A1 and the third thread Post Processor A2, and the types of the third thread Post Processor A1 and the third thread Post Processor A2 are the same. The first thread Data Loader A1, the first thread Data Loader A2, and the first thread Data Loader A3 are executed in a multi-thread concurrent manner; the third thread Post Processor A1 and the third thread Post Processor A2 are executed in a multi-thread concurrent manner; a model operation unit (Model Runner) is created between the first thread Data Loader and the third thread Post Processor, a second cache queue Basic Queue1 is set between the first thread Data Loader and the model operation unit Model Runner, and a third cache queue Basic Queue2 is set between the model operation unit Model Runner and the third thread Post Processor. As shown in the figure, the first thread Data Loader A1, the first thread Data Loader A2, and the first thread Data Loader A3 are the same type of Data Loader, they run in parallel, and the running results of each of them can be put into the second cache queue Basic Queue1 without synchronization. The third thread Post Processor A1 and the third thread Post Processor A2 are the same type of Post Processor, they run in parallel, and each of them independently takes data from the third cache queue Basic Queue2. The specific calculation process is as follows:

[0114] A second cache queue Basic Queue 1 is created and arranged between the first thread Data Loader A1, the first thread Data Loader A2, the first thread Data Loader A3 and the model operation unit Model Runner in parallel arrangement. The results of the parallel calculation of the first thread Data Loader A1, the first thread Data Loader A2, the first thread Data Loader A3 are put into the second cache queue Basic Queue 1. The second cache queue Basic Queue 1 is characterized in that it can start the next execution as soon as a first thread Data Loader pushes data. The model operation unit Model Runner pops one piece of data in the second cache queue Basic Queue 1 at a time. A third cache queue Basic Queue 2 is created and arranged between the model operation unit Model Runner and the third thread Post Processor A1 and the third thread Post Processor A2 in parallel arrangement. The third cache queue Basic Queue 2 is characterized in that the model operation unit Model Runner pushes one piece of data to the third cache queue Basic Queue 2 at a time, and the third thread Post Processor pops one piece of data from the third cache queue Basic Queue 2 at a time. One piece of data includes the model inference result of the first thread Data Loader A1, the model inference result of the first thread Data Loader A2, or the model inference result of the first thread Data Loader A3.

[0115] The running starts, at this time, the second cache queue Basic Queue 1 and the third cache queue Basic Queue 2 are empty, therefore, the model operation unit Model Runner is in the waiting state with the third thread Post Processor A1 and the third thread Post Processor A2. The first thread Data Loader A1, the first thread Data Loader A2 and the first thread Data Loader A3 trigger execution. After the first thread Data Loader A1 has data put into the second cache queue Basic Queue 1, the model operation unit Model Runner is triggered, after the model operation unit Model Runner pops the data, the first thread Data Loader A1 and the model operation unit Model Runner execute concurrently. After the model operation unit Model Runner runs and pushes the data to the third cache queue Basic Queue 2 associated with the third thread Post Processor, the third thread Post Processor triggers running, the model operation unit Model Runner and the third thread Post Processor execute concurrently.

[0116] After the first thread Data Loader A1 processes all data, there is no data in the second cache queue Basic Queue 1 and the third cache queue Basic Queue 2, at this time, the model operation unit Model Runner ends execution, then, the third thread Post Processor also ends running. The inference process ends.

[0117] It needs to be explained that, Figure 9 and Figure 10 The structure of the single pipeline in the embodiment, when the logic of processing data on each first thread is the same, the logic of processing data on each third thread pipeline can be the same or different, Figure 9 and Figure 10 This is only an example of the third thread with the same data processing logic, which is not limited here. When the logic of processing data of the third thread is different, the structure of the single pipeline described in the embodiment can also be used Figure 9 and Figure 10 The specific calculation process is consistent with the calculation process described above, which is not described here.

[0118] When the logic of processing data on multiple first threads is different, the processor performs the following steps when executing the step S301: synchronizing the pre-processing results output by each first thread, and storing the synchronized pre-processing results of each first thread to the second cache queue.

[0119] In this embodiment, after the plurality of first threads perform parallel preprocessing on the to-be-processed data of a model, the preprocessing results of the first threads need to be synchronized, so that the processor simultaneously obtains the preprocessing results of the first threads, and then the synchronized preprocessing results of all the first threads are stored in the second storage queue.

[0120] After the processor stores the preprocessing results of the first threads in the second cache queue, the processor can execute the step S302, and specifically can execute: calling the second thread to load the offline model, the processor can simultaneously input the preprocessing results of all the first threads into the loaded offline model for inference from the second cache queue to obtain the model inference results of the second threads; the processor can input the preprocessing results of all the first threads into the loaded offline model for inference in sequence from the second cache queue to obtain the model inference results of the second threads.

[0121] After the processor obtains the model inference results of the second threads, the processor can execute the step S302, and specifically can execute: copying the model inference results of the second threads according to a preset number of copies to obtain a copy result of the preset number of copies, and storing the copy result of the preset number of copies in the third cache queue.

[0122] The preset number of copies is equal to the number of the third threads, and the preset number of copies can be set by the processor in advance according to the number of the third threads. For example, if there are three third threads, the corresponding preset number of copies is three, that is, the model inference results of the second threads are copied three times.

[0123] In this embodiment, after the processor obtains the model inference results of the second threads, the processor can further copy the model inference results of the second threads according to the preset number of copies to obtain a copy result of the preset number of copies, so that each copy result includes the model inference results of each second thread, and then the copy result of the preset number of copies is stored in the third cache queue according to the arrangement mode of the number of copies.

[0124] After the processor stores the copy result in the third cache queue, the processor can execute the step S303, and specifically can execute: calling each third thread to read a copy result from the third cache queue for parallel post-processing.

[0125] In this embodiment, the number of first threads can be consistent with the number of third threads, or the number of first threads can be inconsistent with the number of third threads. The third cache queue stores multiple copies of results, that is, multiple copies of model inference results of the second threads, and each third thread can read one copy of results from the third cache queue for parallel post-processing. For example, in the case where the number of first threads is consistent with the number of third threads, assuming that there are three first threads, #1 first thread, #2 first thread, and #3 first thread, there are three third threads, #1 third thread, #2 third thread, and #3 third thread, and the third cache queue stores three copies of results, each of which includes the model inference result of the #1 second thread obtained by the #1 first thread after model inference on the second thread, the model inference result of the #2 second thread obtained by the #2 first thread after model inference on the second thread, and the model inference result of the #3 second thread obtained by the #3 first thread after model inference on the second thread, then the #1 third thread reads one copy of results from the third cache queue, the #2 third thread reads one copy of results from the third cache queue, and the #3 third thread reads one copy of results from the third cache queue, and the copies of results read by the #1 third thread, the #2 third thread, and the #3 third thread from the third cache queue are the same.

[0126] For another example, in the case where the number of first threads is inconsistent with the number of third threads, assuming that there are three first threads, #1 first thread, #2 first thread, and #3 first thread, there are two third threads, #1 third thread and #2 third thread, and the third cache queue stores two copies of results, each of which includes the model inference result of the #1 second thread, the model inference result of the #2 second thread, and the model inference result of the #3 second thread, then the #1 third thread reads one copy of results from the third cache queue, the #2 third thread reads one copy of results from the third cache queue, and the copies of results read by the #1 third thread and the #2 third thread from the third cache queue are the same.

[0127] The exemplary description of the present application is in the case where the number of first threads is consistent with the number of third threads, and the above-mentioned data processing method on the pipeline corresponding to the different logic of processing data on each first thread in the above-mentioned embodiment is as follows: Figure 11As shown, it is assumed that there are three first threads, i.e., the first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1, and the types of the first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are different, and there are three third threads, i.e., the third thread Post Processor A1, the third thread Post Processo B1, and the third thread Post Processo C1, and the types of the third thread Post Processor A1, the third thread Post Processo B1, and the third thread Post Processo C1 are different. The first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are executed in a multi-thread concurrent manner; the third thread Post Processor A1, the third thread Post Processo B1, and the third thread Post Processo C1 are executed in a multi-thread concurrent manner; a model operation unit Model Runner is created between the first thread Data Loader and the third thread Post Processor, a second cache queue Sync Queue is created between the first thread Data Loader and the model operation unit Model Runner, and a third cache queue Copy Queue is created between the model operation unit Model Runner and the third thread Post Processor. As shown in FIG. 11, the first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are different types of Data Loaders, they run in parallel, and the running results of each of them need to be synchronized, i.e., put into the second cache queue Sync Queue, and the third thread Post Processor A1, the third thread Post Processo B1, and the third thread Post Processo C1 are different types of Post Processors, they run in parallel, and each of them independently takes data from the third cache queue Copy Queue. The specific calculation process is as follows:

[0128] A second cache queue Sync Queue is created and arranged between the first threads Data Loader A1, Data Loader A2, and Data Loader A3 arranged in parallel and the model operation unit Model Runner. The results of the parallel calculation of the first threads Data Loader A1, Data Loader A2, and Data Loader A3 are placed in the second cache queue Sync Queue. The second cache queue Sync Queue is characterized in that it starts the next execution as soon as a first thread Data Loader pushes data. The model operation unit Model Runner pops one portion of data in the second cache queue Sync Queue at a time. A third cache queue Copy Queue is created and arranged between the model operation unit Model Runner and the third threads Post Processor A1, Post Processor A2, and Post Processor A3 arranged in parallel. The third cache queue Copy Queue is characterized in that the model operation unit Model Runner copies the model inference results of each second thread according to a preset number of copies each time the model inference results of each second thread are output, and pushes the model inference results of each second thread to the third cache queue Copy Queue. The third threads Post Processor pop one portion of data from the third cache queue Copy Queue at a time. One portion of data includes the model inference results of the first thread Data Loader A1, the model inference results of the first thread Data Loader A2, and the model inference results of the first thread Data Loader A3.

[0129] The running starts, at this time, the second cache queue Sync Queue and the third cache queue Copy Queue are empty, therefore, the model operation unit Model Runner and the third threads Post Processor A1, Post Processor B1, Post Processor C1 are in the waiting state. The first threads Data Loader A1, Data Loader B1, Data Loader C1 trigger execution. After the pre-processing threads have data put into the second cache queue Sync Queue, the model operation unit Model Runner is triggered, the model operation unit Model Runner pops all the data, and the first threads Data Loader and the model operation unit Model Runner execute concurrently. After the model operation unit Model Runner runs, the output of each second thread model inference result is copied three times to the third cache queue Copy Queue, and the third threads Post Processor trigger running and each third thread Post processor can pop data in turn. After the last Post Processor pops data, the model operation unit Model Runner and the third threads Post Processor execute concurrently.

[0130] After the first threads Data Loader process all the data, there is no data in the second cache queue Sync Queue and the third cache queue Copy Queue, at this time, the model operation unit Model Runner ends execution, then, the third threads Post Processor also ends running. The inference process ends.

[0131] Optionally, the exemplary description of the present application is in the case where the number of first threads is not consistent with the number of third threads, the above-mentioned logical processing data on each first thread does not correspond to the data processing method on the pipeline, for example, Figure 12As shown, there are three first threads, i.e., first thread Data Loader A1, first thread Data Loader B1, and first thread Data Loader C1, and the types of the first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are different. There are three third threads, i.e., third thread Post Processor A1 and third thread Post Processor B1, and the types of the third thread Post Processor A1 and the third thread Post Processor B1 are different. The first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are executed in a multi-thread concurrent manner; the third thread Post Processor A1 and the third thread Post Processor B1 are executed in a multi-thread concurrent manner; a model operation unit Model Runner is created between the first thread Data Loader and the third thread Post Processor, a second cache queue Sync Queue is created between the first thread Data Loader and the model operation unit Model Runner, and a third cache queue Copy Queue is created between the model operation unit Model Runner and the third thread Post Processor. As shown in FIG. 12, the first thread Data Loader A1, the first thread Data Loader B1, and the first thread Data Loader C1 are different types of Data Loaders, which run in parallel and the running results of each of them need to be synchronized, i.e., put into the second cache queue Sync Queue. The third thread Post Processor A1 and the third thread Post Processor B1 are different types of Post Processors, which run in parallel and independently take data from the third cache queue Copy Queue. The specific calculation process is as follows:

[0132] A second cache queue Sync Queue is created and arranged between the first threads Data Loader A1, Data Loader A2, and Data Loader A3 arranged in parallel and the model operation unit Model Runner. The results of the parallel calculation of the first threads Data Loader A1, Data Loader A2, and Data Loader A3 are put into the second cache queue Sync Queue. The second cache queue Sync Queue is characterized in that it starts the next execution as soon as a first thread Data Loader pushes data. The model operation unit Model Runner pops one portion of data in the second cache queue Sync Queue at a time. A third cache queue Copy Queue is created and arranged between the model operation unit Model Runner and the third threads Post Processor A1 and Post Processor A2 arranged in parallel. The third cache queue Copy Queue is characterized in that the model operation unit Model Runner copies the model inference results of each second thread according to a preset number of copies each time the model inference results of each second thread are output, and pushes the model inference results of each second thread into the third cache queue Copy Queue. The third threads Post Processor pop one portion of data from the third cache queue Copy Queue at a time. One portion of data includes the model inference results of the first thread Data Loader A1, the model inference results of the first thread Data Loader A2, and the model inference results of the first thread Data Loader A3.

[0133] The running starts, at this time, the second cache queue Sync Queue and the third cache queue Copy Queue are empty, therefore, the model operation unit Model Runner and the third thread Post Processor A1 and the third thread Post Processor B1 are in the waiting state. The first thread Data Loader A1, the first thread Data Loader B1 and the first thread Data Loader C1 trigger execution. After the pre-processing thread has data put into the second cache queue Sync Queue, the model operation unit Model Runner is triggered, after the model operation unit Model Runner pops all the data, the first thread Data Loader and the model operation unit Model Runner execute concurrently. After the model operation unit Model Runner runs, the output of each second thread model inference result is copied twice and transmitted to the third cache queue Copy Queue, then the third thread Post Processor triggers running and each third thread Post processor can pop data in turn, after the last PostProcessor pops data, the model operation unit Model Runner and the third thread Post Processor execute concurrently.

[0134] After the first thread Data Loader processes all the data, there is no data in the second cache queue Sync Queue and no data in the third cache queue Copy Queue, at this time, the model operation unit Model Runner ends execution, then the third thread Post Processor also ends running. The inference process ends.

[0135] It should be noted that, Figure 11 and Figure 12 The structure of the single pipeline in the embodiment, wherein when the logic of processing data on each first thread is different, the logic of processing data on each third thread pipeline can be the same or different, Figure 11 and Figure 12 This is only an example of a third thread with different data processing logic, which is not limited here. When the logic of processing data of the third thread is the same, the structure of the single pipeline described in the embodiment can also be used Figure 11 and Figure 12 The specific calculation process is consistent with the calculation process described above, and details are not repeated here.

[0136] In the above embodiments, the same or different types of first threads can be run in parallel through multi-threading, and the different threads can be run in pipeline through the cache queue. The length of the cache queue depends on the number of different types of first threads, that is, the more the number of first threads, the longer the length of the cache queue. The above embodiments simultaneously realize the parallelization within a single pipeline and the pipeline data processing process. When facing the inference of multiple offline models, the inference speed and efficiency can be greatly improved.

[0137] It should be noted that the model inference method for realizing parallelization and streamlining within a single pipeline is described by taking the first pipeline as an example. Of course, the second pipeline can also execute the model inference method according to the above method. The data input into the first pipeline is the model to-be-processed data, and the data input into the second pipeline is the processing result of the first pipeline. The processing or inference method of the data within the second pipeline is described above and will not be described here.

[0138] In combination with all the above embodiments, the present application provides a model inference method for offline models, which can be applied to the offline model inference structure as shown in Figure 13 , which combines Figure 3 and Figure 9The offline model inference structure shown in the embodiment, that is, the model inference method of parallelization and pipelining of multiple same type pipelines, the specific inference process includes: the first thread Data Loader A1, the first thread Data Loader A2, and the first thread Data Loader A3 in the first pipeline A1 are triggered to execute in parallel; the first thread Data Loader A4, the first thread Data Loader A5, and the first thread Data Loader A6 in the first pipeline A2 are triggered to execute in parallel; the first thread Data Loader A7, the first thread Data Loader A8, and the first thread Data Loader A9 in the first pipeline A3 are triggered to execute in parallel. The first pipeline A1, the first pipeline A2, and the first pipeline A3 are triggered to execute in parallel. For each first pipeline, when any first thread in the first pipeline pre-processes the loaded to-be-processed data and puts the pre-processed data into the second cache queue Basic Queue1, the model operation unit Model Runner is triggered, the model operation unit Model Runner reads the pre-processed data from the second cache queue Basic Queue1, and loads the corresponding model to perform inference operation based on the read pre-processed data to obtain the model inference result of the corresponding second thread, and then stores the model inference result of the second thread into the third cache queue Basic Queue2, and the third thread reads the model inference result of the second thread from the third cache queue Basic Queue2 to perform post-processing, and stores the post-processed data to the first cache queue Basic Queue.

[0139] The first thread Data Loader B1, the first thread Data Loader B2 and the first thread Data Loader B3 in the second pipeline B1 are triggered to execute in parallel; the first thread Data Loader B4, the first thread Data Loader B5 and the first thread Data Loader B6 in the first pipeline B2 are triggered to execute in parallel; the first thread Data Loader B7, the first thread Data Loader B8 and the first thread Data Loader B9 in the first pipeline B3 are triggered to execute in parallel. For each second pipeline, when any of the first threads pre-processes the loaded data and stores the pre-processed data in the second cache queue Basic Queue1, the model operation unit Model Runner is triggered, reads the pre-processed data from the second cache queue Basic Queue1, and loads the corresponding model to perform inference operation based on the read pre-processed data to obtain the model inference result of the corresponding second thread, and then stores the model inference result of the second thread in the third cache queue Basic Queue2. The third thread reads the model inference result of the second thread from the third cache queue Basic Queue2 for post-processing, and stores the post-processed data to the first cache queue Basic Queue. The data loaded by the first thread is the post-processed data output by the first pipeline. The above is only a simple description, and the specific implementation of the inference calculation method of the embodiment will be described below with reference to the model inference process shown in the embodiments. Figure 13 The specific inference calculation method of the embodiment can be referred to the model inference process shown in the foregoing Figure 3 and Figure 9 embodiments, which will not be described herein.

[0140] The application also provides an offline model inference method, which can be applied to the offline model inference structure as shown in Figure 14 which combines the Figure 6 and Figure 11The offline model inference structure shown in the embodiment, that is, the offline model inference method which realizes parallelization and pipelining of multiple different types of pipelines, the specific inference process includes: the first threads Data Loader A1, Data Loader B1, Data Loader C1 in the first pipeline Pipeline A1 are triggered to execute in parallel; the first threads Data Loader D1, Data Loader E1, Data Loader F1 in the first pipeline Pipeline B1 are triggered to execute in parallel; the first threads Data Loader G1, Data Loader H1, Data Loader K1 in the first pipeline Pipeline C1 are triggered to execute in parallel. The first pipeline Pipeline A1, the first pipeline Pipeline B1, and the first pipeline Pipeline C1 are triggered to execute in parallel. For each first pipeline, when any of the first threads pre-processes the loaded to-be-processed data and puts the pre-processed data into the second cache queue Sync Queue, the model operation unit Model Runner is triggered, the model operation unit Model Runner reads the pre-processed data from the second cache queue Sync Queue, and loads the corresponding model to perform inference operation based on the read pre-processed data to obtain the model inference result of the corresponding second thread, then the model inference result of the second thread is copied to the third cache queue Copy Queue according to the preset score, the third thread reads out the model inference result of each second thread from the third cache queue Copy Queue for post-processing, and stores the post-processed data to the first cache queue Sync Queue, and copies the post-processed data according to the preset number of copies to obtain the copied data, and stores the copied data to the first copy queue Copy Queue.

[0141] The first thread Data Loader A1, the first thread Data Loader B1 and the first thread Data Loader C1 in the second pipeline Pipeline D1 are triggered to execute in parallel; the first thread Data Loader D1, the first thread Data Loader E1 and the first thread Data Loader F1 in the second pipeline Pipeline E1 are triggered to execute in parallel; the first thread Data Loader G1, the first thread Data Loader H1 and the first thread Data Loader K1 in the second pipeline Pipeline F1 are triggered to execute in parallel. For each second pipeline, when any of the first threads pre-processes the loaded data and puts the pre-processed data into the second cache queue Sync Queue, the model operation unit Model Runner is triggered, the model operation unit Model Runner reads the pre-processed data from the second cache queue Sync Queue, and loads a corresponding model to perform inference operation based on the read pre-processed data to obtain a model inference result of a corresponding second thread, and then copies the model inference result of the second thread to the third cache queue Copy Queue according to a preset number of copies, and the third thread reads the model inference result of the second thread from the third cache queue Copy Queue to perform post-processing, and stores the post-processed data to the first cache queue Basic Queue. The data loaded by the first thread is the copied data stored in the first copy queue Copy Queue. The above is only a simple description, and the specific inference calculation method of the embodiment will be described below with reference to the model inference process shown in the embodiments. Figure 14 The specific inference calculation method of the embodiment can be referred to the model inference process shown in the foregoing Figure 6 and Figure 11 The model structure described in the embodiment is only an application description,

[0142] It should be noted that, Figure 13 and Figure 14 The model structure described in the embodiment is only an application description, Figure 13 The types of the first pipelines are the same, the types of the second pipelines are the same, the logic of processing data of the first threads in each first pipeline is the same, and the logic of processing data of the first threads in each second pipeline is the same. Figure 14 The types of the first pipelines are different, the types of the second pipelines are different, the logic of processing data of the first threads in each first pipeline is different, and the logic of processing data of the first threads in each second pipeline is different. In actual application, Figure 13 and Figure 14In the offline model inference structure, the types of the first pipelines can be the same, and the types of the second pipelines are different; or the types of the first pipelines can be different, and the types of the second pipelines are the same; or the types of the first pipelines can be different, and the types of the second pipelines are different. For the structure in each pipeline, the logic of the first thread processing data in each pipeline can be the same, and the logic of the third thread processing data is the same; optionally, the logic of the first thread processing data in each pipeline can be the same, and the logic of the third thread processing data is different; optionally, the logic of the first thread processing data can be different, and the logic of the third thread processing data is the same; optionally, the logic of the first thread processing data can be different, and the logic of the third thread processing data is different. Correspondingly, Figure 13 and Figure 14 In the offline model inference structure, the number of the first pipelines and the number of the second pipelines are the same, and the number of the first threads and the number of the third threads included in each pipeline are the same. However, in actual application, the number of the first pipelines and the number of the second pipelines can be the same, and the number of the first threads and the number of the third threads included in each pipeline are not the same; optionally, the number of the first pipelines and the number of the second pipelines can be different, and the number of the first threads and the number of the third threads included in each pipeline are the same; optionally, the number of the first pipelines and the number of the second pipelines can be different, and the number of the first threads and the number of the third threads included in each pipeline are not the same. The type of each pipeline and the logic of the threads processing data in each pipeline can be determined according to actual inference requirements.

[0143] It should be understood that, although Figure 2 , Figure 5 and Figure 8 the steps in the flowcharts are shown in sequence according to the arrows, these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, Figure 2 , Figure 5 and Figure 8 At least part of the steps can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or steps or stages in other steps.

[0144] In one embodiment, as shown in Figure 15 , an offline model inference apparatus is provided, comprising: an acquisition module 11, a first parallel processing module 12 and a second parallel processing module 13, wherein:

[0145] The acquisition module 11 is used to acquire the data to be processed from multiple offline models.

[0146] The first parallel processing module 12 is used to perform first parallel processing on the data to be processed of the multiple offline models on multiple first pipelines, and to store the processing results output by each first pipeline to a first buffer queue.

[0147] The second parallel processing module 13 is used to read the processing results output by each of the first pipelines from the first buffer queue on multiple second pipelines and perform second parallel processing to complete the inference of the multiple offline models.

[0148] In one embodiment, if the multiple offline models are of the same type, the first parallel processing module 12 is specifically used to store the processing results output by each of the first pipelines into the first cache queue in the order in which the processing results are output by each of the first pipelines.

[0149] In one embodiment, the second parallel processing module 13 is specifically used to read at least one processing result output by the first pipeline from the first buffer queue on each of the second pipelines and perform second parallel processing.

[0150] In one embodiment, if the plurality of models are offline models of different types, the first cache queue includes a first storage queue and a first copy queue, such as... Figure 16 As shown, the first parallel processing module 12 includes:

[0151] The synchronization unit 121 is used to synchronize the processing results output by each of the first pipelines and store the synchronized processing results output by each of the first pipelines into the first storage queue.

[0152] The copy unit 122 is used to copy the processing results output by each of the first pipelines in the first storage queue according to a preset number of copies, so as to obtain the preset number of copy results, and store the preset number of copy results in the first copy queue, wherein the preset number of copies is equal to the number of the second pipelines.

[0153] In one embodiment, the second parallel processing module 13 is specifically used to read one copy result from the first copy queue on each of the second pipelines and perform second parallel processing.

[0154] In one embodiment, each of the first pipelines is executed using multiple threads, including multiple first threads, one second thread, and multiple third threads, such as... Figure 17 As shown, the first parallel processing module 12 includes:

[0155] The pre-processing unit 123 is configured to invoke the plurality of first threads to perform parallel pre-processing on the data to be processed of the offline model corresponding to each of the first pipelines, and store the pre-processing results of each of the first threads to a second cache queue.

[0156] The inference unit 124 is configured to invoke the second threads to load the model and perform model inference based on the pre-processing results in the second cache queue, obtain model inference results of each of the second threads, and store the model inference results of each of the second threads to a third cache queue.

[0157] The post-processing unit 125 is configured to invoke the plurality of third threads to read the model inference results of each of the second threads from the third cache queue for parallel post-processing.

[0158] In an embodiment, the plurality of first threads and the plurality of third threads run on a first processor, and the second threads run on a second processor.

[0159] In an embodiment, if the logic of processing data on the plurality of first threads is the same, the pre-processing unit 123 is specifically configured to store the pre-processing results of each of the first threads to the second cache queue in sequence according to the order of outputting the pre-processing results of each of the first threads.

[0160] In an embodiment, the post-processing unit 125 is specifically configured to invoke each of the third threads to read at least one of the model inference results of the second threads from the third cache queue for parallel post-processing.

[0161] In an embodiment, if the logic of processing data on the plurality of first threads is different, the pre-processing unit 123 is specifically configured to synchronize the pre-processing results output by each of the first threads, and store the synchronized pre-processing results of each of the first threads to the second cache queue.

[0162] In an embodiment, the copying unit 122 is specifically configured to copy the model inference results of each of the second threads according to a preset number of copies to obtain a preset number of copy results, and store the preset number of copy results to the third cache queue; the preset number of copies is equal to the number of the third threads.

[0163] In an embodiment, the post-processing unit 125 is specifically configured to invoke each of the third threads to read one of the copy results from the third cache queue for parallel post-processing.

[0164] The specific limitations of the device for offline model inference can be referred to the limitations of the method for offline model inference in the above, which will not be repeated here. Each module in the device for offline model inference described above can be realized by software, hardware and their combinations in whole or in part. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so that the processor calls and executes the operations corresponding to each module. In one embodiment, a computer device is provided, comprising a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the following steps:

[0165] Obtaining the to-be-processed data of the plurality of offline models;

[0166] Respectively performing first parallel processing on the to-be-processed data of the plurality of offline models on a plurality of first pipelines, and storing the processing results output by each of the first pipelines to a first cache queue;

[0167] Respectively reading the processing results output by each of the first pipelines from the first cache queue to perform second parallel processing on a plurality of second pipelines to complete the inference of the plurality of offline models.

[0168] Depending on the application scenario, the computing devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, camcorders, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The processors or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the processors or apparatus disclosed herein can also be used in cloud, edge, and terminal applications related to artificial intelligence, big data, and / or cloud computing. In one or more embodiments, the high-performance processors or devices according to this disclosure can be applied to cloud devices (e.g., cloud servers), while low-power processors or devices can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration. The computer device provided in the above embodiments has a similar implementation principle and technical effect to the above method embodiments, and will not be repeated here.

[0169] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0170] Obtain the data to be processed from multiple offline models;

[0171] The data to be processed from the multiple offline models are processed in parallel on multiple first pipelines, and the processing results output from each first pipeline are stored in a first buffer queue.

[0172] The processing results output from each of the first pipelines are read from the first buffer queue on multiple second pipelines and then processed in a second parallel manner to complete the inference of the multiple offline models.

[0173] The computer readable storage medium provided by the above embodiment has similar implementation principles and technical effects to the method embodiments, and thus will not be described here.

[0174] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, the RAM can be static random access memory (SRAM) or dynamic random access memory (DRAM).

[0175] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.

[0176] The above embodiments only express several implementation manners of the present application, and the description is specific and detailed, but it should not be considered as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of the patent protection of the present application should be subject to the appended claims.

Claims

1. A method for offline model inference, characterized in that, The method employs a multi-stage pipeline, which includes multiple first pipelines and multiple second pipelines. The method includes: Obtain the data to be processed from multiple offline models; The data to be processed from the multiple offline models are processed in parallel on multiple first pipelines, and the processing results output from each first pipeline are stored in a first buffer queue. The processing results output from each of the first pipelines are read from the first buffer queue on multiple second pipelines and then processed in a second parallel manner to complete the inference of the multiple offline models. Each of the pipelines is executed using multiple threads, including multiple first threads, one second thread, and multiple third threads. The multiple first threads and the multiple third threads run on a first processor, and the second thread runs on a second processor.

2. The method according to claim 1, characterized in that, If the multiple offline models are of the same type, storing the processing results output by each of the first pipelines into the first cache queue includes: The processing results output by each of the first pipelines are sequentially stored in the first buffer queue according to the order in which they are output.

3. The method according to claim 2, characterized in that, The step of reading the processing results output from the first pipeline from the first buffer queue on multiple second pipelines and performing second parallel processing includes: In each of the second pipelines, at least one processing result output from the first pipeline is read from the first buffer queue and processed in the second parallel process.

4. The method according to claim 1, characterized in that, If the multiple offline models are of different types, the first cache queue includes a first storage queue and a first copy queue. Storing the processing results output from each of the first pipelines into the first cache queue includes: Synchronize the processing results output by each of the first pipelines, and store the synchronized processing results output by each of the first pipelines into the first storage queue; The processing results output by each of the first pipelines in the first storage queue are copied according to a preset number of copies to obtain the preset number of copy results, and the preset number of copy results are stored in the first copy queue, wherein the preset number of copies is equal to the number of the second pipelines.

5. The method according to claim 4, characterized in that, The step of reading the processing results output from the first pipeline from the first buffer queue on multiple second pipelines and performing second parallel processing includes: On each of the second pipelines, a copy result is read from the first copy queue and processed in the second parallel process.

6. The method according to claim 1, characterized in that, The method includes: For each pipeline, the plurality of first threads are invoked to perform parallel preprocessing on the data to be processed of the offline model corresponding to the pipeline, and the preprocessing results of each first thread are stored in the second cache queue. The second thread is invoked to load the offline model, and the offline model is inferred based on the preprocessing results in the second cache queue to obtain the model inference results of the second thread, and the model inference results of the second thread are stored in the third cache queue. The multiple third threads are invoked to read the model inference results of each second thread from the third cache queue and perform parallel post-processing.

7. The method according to claim 6, characterized in that, If the logic for processing data is the same on the multiple first threads, then storing the preprocessing results of each first thread into the second cache queue includes: The preprocessing results of each of the first threads are sequentially stored in the second cache queue according to the order in which they output the preprocessing results.

8. The method according to claim 7, characterized in that, The step of calling the multiple third threads to read the model inference results of each of the second threads from the third cache queue and performing parallel post-processing includes: Each of the third threads is invoked to read at least one model inference result from the second thread from the third cache queue and perform parallel post-processing.

9. The method according to claim 6, characterized in that, If the logic for processing data differs across the multiple first threads, then the step of calling the second thread to store the preprocessing results of each first thread into the second cache queue includes: The preprocessing results output by each of the first threads are synchronized, and the synchronized preprocessing results of each of the first threads are stored in the second cache queue.

10. The method according to claim 9, characterized in that, The step of storing the model inference results of each of the second threads into the third cache queue includes: The model inference results of each of the second threads are copied according to a preset number of copies to obtain the preset number of copy results, and the preset number of copy results are stored in the third cache queue; the preset number of copies is equal to the number of the third threads.

11. The method according to claim 10, characterized in that, The process involves invoking multiple third threads to read the model inference results of each second thread from the third cache queue and performing parallel post-processing, including: Each of the third threads is invoked to read a copy of the result from the third cache queue and perform parallel post-processing.

12. An apparatus for offline model inference, characterized in that, The device includes: The acquisition module is used to acquire data to be processed from multiple offline models; The first parallel processing module is used to perform first parallel processing on the data to be processed of the multiple offline models on multiple first pipelines respectively, and store the processing results output by each first pipeline to the first cache queue. The second parallel processing module is used to read the processing results output by each of the first pipelines from the first buffer queue in multiple second pipelines and perform second parallel processing to complete the inference of the multiple offline models. Each of the first pipelines is executed using multiple threads, including multiple first threads, one second thread, and multiple third threads. The multiple first threads and the multiple third threads run on a first processor, and the second thread runs on a second processor.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method for visually designing assembly line and readable storage medium

    CN112667227A

  • Multi-model parallel reasoning method based on AI chip

    CN112783650A