Artificial intelligence application deployment method based on domestic GPU

By using the OpenCL framework to compute kernels in parallel on domestic GPUs, the problem of insufficient computing performance of domestic GPUs is solved, and high-performance deployment and real-time improvement of AI applications are achieved.

CN120353475APending Publication Date: 2025-07-22CHINESE AERONAUTICAL RADIO ELECTRONICS RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510427989.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the computing performance of domestic GPUs is weaker than that of foreign GPUs, resulting in the need to transmit data through the network when deploying AI applications, which increases system delay and is not conducive to independent control and domestic production.

Method used

Using the OpenCL programming framework to deploy AI applications on domestic GPUs, analyze parallel computing kernels, and use the computing power of domestic GPUs to perform local processing, including parallel acceleration of operators such as convolution, pooling, activation, full connection and LSTM.

Benefits of technology

It realizes high-performance deployment of AI applications on domestic GPUs, reduces system delay, reduces dependence on network transmission, and improves real-time and autonomous controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353475A_ABST
    Figure CN120353475A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence application deployment method based on a domestic GPU (Graphics Processing Unit), which comprises the following steps: loading an algorithm model of an artificial intelligence application, analyzing and extracting neural network operators used in the algorithm model, and forming a calculation flow graph according to an operator sequence in an original algorithm model; the method comprises the following steps: selecting domestic GPU equipment from an OpenCL platform, creating a context based on the GPU equipment, creating a command queue on the context, and creating a kernel object and a memory object for each operator by utilizing a cache region object; writing a parallel computing kernel of each operator by using OpenCL, and generating a program object and an instantiated kernel of the operators at the same time; and sequentially calling the kernel according to a front-back connection relationship of operators in the calculation flow graph, and operating the kernel on the domestic GPU equipment, thereby completing deployment of the artificial intelligence application. According to the invention, the AI application can be deployed with high performance, network transmission is not needed, and the time delay is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embedded deployment of high-performance computing and artificial intelligence applications based on domestic GPUs, and particularly relates to a method for deploying artificial intelligence applications based on domestic GPUs. Background Art

[0002] With the rapid rise of artificial intelligence technology (AI), people's life, work and study have undergone earth-shaking changes. In an autonomous driving system, AI technology can be used to perceive the environment, identify objects, fuse sensor data, and plan paths, thereby assisting people in "driving" and liberating people's minds. In face recognition and speech recognition translation systems, AI technology enables most companies to quickly manage and count the attendance of relevant personnel for work, and can also perform secure unlocking of mobile phones and autonomous translation of other foreign languages, greatly improving the office efficiency of the company and solving the problem of language communication in multiple regions. In the military field, such as in intelligence, surveillance and reconnaissance systems, unmanned aerial vehicle swarms combined with AI technology can be used for target recognition and detection, thereby predicting and judging the trajectories of the enemy, and finally effectively damaging and attacking the targets; in the equipment health management system, AI technology can check the integrity of air and ground vehicles before and after battlefield operations, intelligently predict the possible time of equipment failures, and allow problems to be solved before the mission is affected, thereby improving the safety of soldiers and equipment.

[0003] Ultimately, the implementation of the above-mentioned AI applications is inseparable from the support of big data, hardware computing devices and software computing frameworks. When deploying AI applications, it is usually necessary to transmit big data through the network to a cloud server for online inference of the above-mentioned AI technology, and finally feedback the processed results to the local terminal. In this process, both the front-end sensor data and the intermediate inference data generated need to be transmitted through the network as an intermediate device, which brings unnecessary delays to the entire system. In addition, since the current hardware devices are mainly from foreign companies such as Nvidia and Intel, it is not conducive to independent control and localization. At present, the domestic hardware device platforms mainly rely on CPUs and GPUs. CPUs can provide complex logical judgment and processing. Compared with CPUs, GPUs can provide greater computing power, but their computing performance is still ten times weaker than that of foreign counterparts.

[0004] In view of the characteristics of domestic GPU platforms, which have low power consumption but relatively weaker computing performance compared to foreign GPUs, there is an urgent need to build a framework on them that can deploy AI applications. By using OpenCL (Open Computing Language) to write parallel computing kernels, the computing power of domestic GPUs can be fully utilized to achieve local processing of data by AI, thereby getting rid of the dependence on network transmission for data, providing a powerful way for the deployment of AI applications such as target real-time detection, precise identification and positioning, and rapid auxiliary decision-making, and at the same time providing a solution for cross-platform transplantation of general parallel AI computing programs for domestic GPUs. Summary of the Invention

[0005] The object of the present invention is to provide an artificial intelligence application deployment method based on domestic GPUs. First, neural network operators used in AI applications such as target real-time detection, precise identification and positioning, and rapid auxiliary decision-making are parsed. Subsequently, convolutional operators, pooling operators, activation operators, and basic mathematical operation operators used in target real-time detection and precise identification and positioning, and fully connected operators and LSTM operators used in rapid auxiliary decision-making are parallelly accelerated using the unified programming framework language OpenCL. Finally, each operator is run on the domestic GPU platform to achieve the deployment of the entire AI application on domestic GPUs. All computing operators are implemented and encapsulated using C language and the OpenCL standard of parallel language. The overall interface is unified, the operators are efficiently implemented, the computing power of the domestic GPU platform is fully utilized, and the processing effect is good.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] An artificial intelligence application deployment method based on domestic GPUs includes the following steps:

[0008] Step 1: Obtain the data source of the local artificial intelligence application;

[0009] Step 2: Load the algorithm model of the artificial intelligence application, parse and extract the neural network operators used in the algorithm model, and form a computational flow graph according to the operator order in the original algorithm model;

[0010] Step 3: Select a domestic GPU device from the OpenCL platform, create a context clContext based on this GPU device, create a command queue clCommandQueue on the clContext, and at the same time create a kernel object and a memory object for each operator using the buffer object clCreateBuffer;

[0011] Step 4: Use OpenCL to write the parallel computing kernels of each operator, and at the same time generate the program objects and instantiated kernels of these operators;

[0012] Step 5: Calculate the connection relationship before and after the operators in the flow graph according to Step 2, sequentially call the kernels in Step 4, and run the kernels on domestic GPU devices, so as to complete the deployment of the artificial intelligence application and process the data source in Step 1.

[0013] Preferably, the flow graph calculated in Step 2 is a tree structure in the form of a multi-level directory, and each operator exists in the directory file in the form of a key-value pair <index node of the operator, weight data of the operator>.

[0014] Preferably, in Step 4, when implementing the kernel of the convolution operator using OpenCL, since the convolution operation can be regarded as the dot product of the input matrix and the convolution kernel at different positions, this solution converts the convolution operation into a matrix-matrix multiplication algorithm (GEMM) to accelerate its computing efficiency on domestic GPUs. First, all possible convolution windows of the input matrix are extracted to form a two-dimensional matrix, and each row corresponds to the expansion of a convolution window. Similarly, the convolution kernel is also expanded into a matrix, and each column corresponds to the expansion of a convolution kernel. Finally, the convolution operation can be completed through matrix multiplication. The steps are as follows:

[0015] Specifically, assume that the size of the input image X is H×W×C, where H is the height of the input image, W is the width of the input image, and the number of input channels is C; the size of the convolution kernel F is K×K×C out , K is the size of the convolution kernel, and the number is C out ; the stride when sliding during convolution is S; the size of the matrix Y output in the middle process is (H out ×W out , 1), where H out is the height of the output, and W out is the width of the output, The final convolution calculation result matrix Y conv_out has a size of (H out , W out ).

[0016] First, perform a linear transformation on the input image X to convert it into a matrix X′. After the transformation, the size of X′ is (H out ×W out , K×K×C), and each row corresponds to the expansion of a convolution window. The convolution F is converted into a matrix F′ with a size of (K×K×C,1), and each column corresponds to the expansion of a convolution kernel. Finally, the convolution operation can be expressed as a two-dimensional matrix multiplication: Y = X′×F′. The convolution algorithm can be quickly completed through two-dimensional matrix multiplication.

[0017] Preferably, in Step 4, when implementing the kernels of the pooling and activation operators using OpenCL, the steps are as follows:

[0018] Open up the optimal number of threads T for current domestic GPU devices, and the global number of threads is (H * W) / (T * 4);

[0019] When it is an activation operator, the following parallel program is used for calculation, where n is the index of the number of threads:

[0020] output[n] = input[n] > 0? input[n] : input[n] * 0.1

[0021] When it is a pooling operator, the following parallel program is used for calculation:

[0022] output[n] = max(input[n - 3 : n + 3]).

[0023] Preferably, in step 4, when implementing the kernel of the fully connected operator using OpenCL, the steps are as follows:

[0024] Place vector B in the local memory of the GPU with faster memory access, use OpenCL to open up M parallel threads, perform the multiplication calculation of array A and vector B on the M threads, and use shared memory on the workgroup to adopt a reduction algorithm for parallel acceleration of the fully connected operator.

[0025] Preferably, in step 4, when implementing the kernel of the LSTM operator using OpenCL, the steps are as follows:

[0026] (1) Calculate the forget gate and the input gate concurrently, where the forget gate f t is determined by the hidden layer state h t-1 at the previous moment t and the input x

[0027] f t = σ(W f h t-1 + U f x t + b f )

[0028] where W f , U f and b f are the weight and bias matrices related to the forget gate, and σ is the sigmod function;

[0029] The input gate consists of the following two parts:

[0030] i t = σ(W i h t-1 + U i x t + b i )

[0031] a t = σ(W a h t-1 + U a x t + b a )

[0032] where W i , W a , U i , U a , b i and b a are the weight and bias matrices related to the input gate;

[0033] (2) Calculate the update gate:

[0034] C t = C t-1 ⊙ f t + i t ⊙ a t

[0035] where ⊙ is the Hadamard product;

[0036] (3) The output of the output gate is:

[0037] O t = σ(W O h t-1 + U O x t + b O ).

[0038] Preferably, in step 4, when implementing the kernel of the convolution operator using OpenCL, the steps are as follows:

[0039] Allocate a global thread count of H * W and perform the calculation using the following parallel program, where n is the index of the thread count:

[0040] output[n] = input0[n] + input1[n];

[0041] output[n] = input0[n] - input1[n];

[0042] output[n] = input0[n] * input1[n];

[0043] output[n] = input0[n] / input1[n].

[0044] The beneficial effects of the present invention are as follows:

[0045] 1. An AI application deployment method based on domestic GPUs provided by the present invention accelerates operators such as convolution, pooling, activation, fully connected, LSTM, and basic mathematical operations using OpenCL single-instruction multiple-threading on domestic GPUs, enabling high-performance deployment of AI applications.

[0046] 2. An AI application deployment method based on domestic GPUs provided by the present invention can directly utilize source data to deploy AI applications on existing domestic GPU platforms. It does not require transmitting data over the network to a server for processing, nor does it require transmitting the detection results from the server back over the network. Instead, it can directly process data on the domestic GPU platform, thereby reducing system latency and enabling the deployment of AI applications on domestic GPUs. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the overall block diagram of an AI application deployment method based on domestic GPUs shown in the embodiments.

[0048] Figure 2 is the operator computation flow diagram corresponding to the YOLOV5 object detection model in the embodiments.

[0049] Figure 3 is the detection result of directly deploying the YOLOV5 object detection application on a domestic GPU in the embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0051] See Figure 1 As shown, an AI application deployment method based on domestic GPUs shown in this embodiment mainly consists of an AI application layer, an AI model parsing layer, an operator representation layer, an operator high-performance implementation layer, and a hardware layer. Among them, the operator high-performance implementation layer uses the standard OpenCL general computing framework to accelerate operators such as convolution, pooling, activation, fully connected, LSTM, and basic mathematical operations, and can support any domestic GPU device, with strong code generality and portability. At the same time, it directly utilizes local data for high-performance deployment of AI applications without the participation of network data transmission, making the entire system simple and efficient.

[0052] An AI application deployment method based on domestic GPUs includes the following steps:

[0053] Step 1: Obtain the data source of the local AI application.

[0054] For example, in this embodiment, a local video is used to simulate the data source at the sensor source end, and this is used as the excitation source for AI applications such as target real-time detection, precise identification and positioning, and rapid auxiliary decision-making, to detect the AI algorithm model, so that it is not necessary to transmit the data through the network to the server for processing, reducing the latency introduced by network transmission.

[0055] Step 2: Load the algorithm model of the AI application, parse and extract the neural network operators used in the algorithm model. For operators such as convolutional operators, pooling operators, activation operators, fully connected operators, LSTM, and basic mathematical operations in the model, a computational flow graph is formed according to the operator order in the original AI model.

[0056] The algorithm models adopted by different AI applications are different, and the connection order and weight data of operators such as convolution, pooling, activation, fully connected, LSTM, and basic mathematical operations used in the model are also different. Therefore, it is necessary to read the AI algorithm model file, parse the flow information of the operators in the current algorithm model, and convert it into a tree structure in the form of a multi-level directory. In this way, the entire algorithm model takes the operator as the basic unit, and can exist in the directory file in the form of <index node of the operator, weight data of the operator>, thus forming a simple tree-shaped data warehouse. Based on this tree-shaped data warehouse, various operator objects in the algorithm model can be established, and finally the connection relationship and calculation process of the operators are generated through the string pool.

[0057] Step 3: Parallelly accelerate the processing of all the operators in Step 2 on the domestic GPU using the OpenCL language. Obtain the OpenCL platform, select the domestic OpenCL GPU device from the platform, create a context clContext based on this device, create a command queue clCommandQueue on the clContext, and at the same time use the buffer object clCreateBuffer to create kernel objects and memory objects for operators such as convolutional, pooling, activation, fully connected, LSTM, and basic mathematical operations, and implement unified data encapsulation for each operator, providing a standard function interface.

[0058] Step 4: Write the parallel computing kernels for OpenCL operators such as convolutional, pooling, activation, fully connected, LSTM, and basic mathematical operations, and at the same time generate the program objects and instantiated kernels of these operators.

[0059] (a) When implementing a high-performance convolutional operator using OpenCL, using computationally intensive matrix multiplication to replace the corresponding element multiplication between block matrices, combined with the cache characteristics of the domestic GPU hardware, the present invention is implemented using the following steps:

[0060] (1) Perform a linear transformation on the input image X with the arrangement format of H×W×C in memory to convert it into another matrix X′. The specific linear transformation process is as follows: for each output position (i, j) (where i ∈ [0, H out -1], j ∈ [0, W out -1]), the position of the corresponding convolution window in the input matrix is (i×S, j×S), the size of the window is K×K, then the index range of the covered input matrix is from (i×S, j×S) to (i×S + K - 1, j×S + K - 1). Expand the elements of this window into a row, corresponding to the (i×W out +j)-th row in X′. The overall operation can be expressed as:

[0061] X′ (i×Wout+j),l = X (i×S+m),(j×S+n),p

[0062] where l = m×K×C + n×C + p (m ∈ [0, K - 1], n ∈ [0, K - 1], p ∈ [0, C - 1]).

[0063] (2) After the above transformation, the input image X is transformed into X′, and the transformed shape is (H out ×W out , K×K×C). The size of the convolution kernel matrix F′ is (K×K×C, 1), then the matrix multiplication can be expressed as:

[0064] Y = X′×F′

[0065] where Y is the output matrix, with a size of (H out ×W out , 1), and the calculation formula is as follows:

[0066]

[0067] (3) Utilize the powerful parallel computing ability of the GPU to perform C out matrix multiplications. Since the C out matrix multiplications are independent of each other and there is no data dependency, the present invention utilizes the out-of-order feature of the OpenCL command queue to simultaneously and parallelly process the C out matrix multiplications, thereby improving the utilization rate of the OpenCL execution units on domestic GPUs and accelerating the parallel multiplication processing of the convolution operator.

[0068] (4) Output the final convolution calculation result matrix Y through one-dimensional linear index inverse transformation conv_out , with a size of (H out , W out ), and the inverse transformation calculation formula is as follows:

[0069] Y conv_out,i,j = Yi×Wout+j

[0070] (b) When implementing high-performance pooling and activation operators using OpenCL, since such operators store data continuously during calculation, vectorized float4 can be used for direct acceleration. In specific implementation, the number of local threads can be set to the optimal number of threads T for the current domestic GPU device, and the number of global threads is (H * W) / (T * 4). When it is an activation operator, the following parallel program is used for calculation, where n is the index of the number of threads:

[0071] output[n] = input[n] > 0? input[n] : input[n] * 0.1

[0072] When it is a pooling operator, the following parallel program is used for calculation:

[0073] output[n] = max(input[n - 3:n + 3])

[0074] (c) When implementing high-performance fully connected operators using OpenCL, this operator is essentially a dot product operation of array A and vector B, outputting vector O. That is, O(N * 1) = A(N * M) * B(M * 1), where the size of the output O vector is N rows and 1 column, the size of the input A array is N rows and M columns, and the size of the input B vector is M rows and 1 column. It can be seen that there is a lot of data reuse in this operator, and B is fixed-length reduced. Therefore, vector B can be placed in the local memory of the GPU with faster memory access, and M parallel threads are opened using OpenCL to perform multiplication calculations on the M threads, and a reduction algorithm is used on the workgroup using shared memory for parallel acceleration of the fully connected operator.

[0075] (d) When implementing high-performance LSTM operators using OpenCL, since the LSTM operator is essentially the calculation of 4 gated units, namely the forget gate, input gate, update gate, and output gate, and the calculations of the forget gate and input gate are independent of each other without dependency, the out-of-order first concurrent execution of the forget gate and input gate is used using OpenCL, and finally the update gate and output gate are executed sequentially. The specific implementation of the LSTM operator of the present invention is shown in the following steps:

[0076] (1) Concurrent calculation of the forget gate and input gate, where the forget gate f t is determined by the hidden layer state h t-1 at the previous moment t and the input x

[0077] f t = σ(W f h t-1 + U f x t + b f)

[0078] Among them, W f , U f and b f are the weight and bias matrices related to the forget gate, and σ is the sigmoid function. The input gate consists of the following two parts:

[0079] i t = σ(W i h t-1 + U i x t + b i )

[0080] a t = σ(W a h t-1 + U a x t + b a )

[0081] Among them, W i , W a , U i , U a , b i and b a are the weight and bias matrices related to the input gate.

[0082] (2) Calculate the update gate. The update of the cell state C t in the LSTM consists of two parts: the input gate and the forget gate:

[0083] C t = C t-1 ⊙ f t + i t ⊙ a t

[0084] Among them, ⊙ is the Hadamard product, and C t-1 is the cell state output at the previous moment.

[0085] (3) Finally, the output of the output gate in the LSTM is:

[0086] O t = σ(W O h t-1 + U O x t + b O )

[0087] Among them, W o , U o , b o are the weight and bias matrices related to the output gate.

[0088] (e) When implementing high-performance basic mathematical operation operators using OpenCL, such operators are essentially the addition, subtraction, multiplication, and division operations of two operands of the same scale. Therefore, the same number of threads can be created for calculation. Specifically, when implementing, the global number of threads can be set to H*W, and the following parallel program can be used for calculation, where n is the index of the number of threads:

[0089] output[n] = input0[n] + input1[n];

[0090] output[n] = input0[n] - input1[n];

[0091] output[n] = input0[n] * input1[n];

[0092] output[n] = input0[n] / input1[n];

[0093] By using the OpenCL implementation in the above invention, neural network operators such as convolution, pooling, activation, fully connected, LSTM, and basic mathematical operations can be accelerated, enabling artificial intelligence applications to be deployed with high performance on domestic GPUs.

[0094] In Step 3 and Step 4, standard OpenCL kernel programs are established for each operator such as convolution, pooling, activation, fully connected, LSTM, and basic mathematical operations in the domestic GPU, and workgroups and work items that can be parallelized are created for each kernel program in the domestic GPU device. Each work item acts as an independent thread, without affecting each other, and can concurrently execute to achieve the high-performance implementation of the above operators.

[0095] Step 5: Calculate the connection relationship between the operators in the flow graph according to Step 2, sequentially call the OpenCL kernel in Step 4, and run the kernel on the domestic GPU device, thereby completing the deployment of the AI application, processing the data source in Step 1, and finally, the running target detection results can be viewed in real time on the GPU device, without the need to transmit the detection results back to the user through the network from the AI server, reducing the latency introduced by network transmission and improving the real-time performance of the AI application.

[0096] To verify the efficacy of an artificial intelligence application deployment method based on a domestic GPU provided in this implementation, in this example, real-time target detection is performed on a 1080P video using the YOLOV5 AI detection model on a domestic Jingjiawei GPU, and the detection results are displayed for experiments. Figure 2 Calculate the flow graph for the connection relationship of the operators in the YOLOV5 neural network model after framework parsing; Figure 3The image directly displayed after using the deployment framework for localization can be observed. It can be found that people, vehicles, cats, dogs, etc. in the image can be correctly detected and recognized. At the same time, when comparing the time-consuming process where the comparison data needs to be transmitted over the network to the AI server, and then the AI server sends the operation result back to the local for display, it is found that the artificial intelligence application deployment framework based on domestic GPUs can complete the AI target detection application deployment of 1 1080P image directly in local inference and display within 8 ms. The system delay is small, the display is smooth, and the GPU power consumption is only 20W. However, the time required for the network data transmission method is 15 ms, the system delay is large, and the power consumption of the AI server is as high as 75W.

[0097] It can be understood that for those of ordinary skill in the art, equivalent substitutions or changes can be made according to the technical solution of the present invention and its inventive concept, and all such changes or substitutions should fall within the protection scope of the claims appended to the present invention.

Claims

1. An artificial intelligence application deployment method based on domestic GPUs, characterized in that It includes the following steps: Step 1: Obtain the data source of the local artificial intelligence application; Step 2: Load the algorithm model of the artificial intelligence application, parse and extract the neural network operators used in the algorithm model, and form a computational flow graph according to the operator order in the original algorithm model; Step 3: Select a domestic GPU device from the OpenCL platform, create a context clContext based on this GPU device, create a command queue clCommandQueue on the clContext, and at the same time create a kernel object and a memory object for each operator using the buffer object clCreateBuffer; Step 4: Write the parallel computing kernels of each operator using OpenCL, and at the same time generate the program objects and instantiate the kernels of these operators; Step 5: According to the front and back connection relationship of the operators in the computational flow graph in Step 2, sequentially call the kernels in Step 4 and run the kernels on the domestic GPU device, thereby completing the deployment of the artificial intelligence application and processing the data source in Step 1.

2. The method for deploying an artificial intelligence application based on a domestic GPU according to claim 1, wherein The computational flow graph in Step 2 is a tree structure in the form of a multi-level directory, and each operator exists in the directory file in the form of a key-value pair <operator index node, operator weight data>.

3. The artificial intelligence application deployment method based on domestic GPUs according to claim 1, wherein In Step 4, when implementing the kernel of the convolution operator using OpenCL, the steps are as follows: Let the size of the input image X be H×W×C, where H is the height of the input image, W is the width of the input image, and the number of input channels is C; the size of the convolutional kernel F is K×K×C out , where K is the size of the convolutional kernel and the number is C out ; the stride for sliding during convolution is S; the size of the matrix Y output in the intermediate process is (H out ×W out , 1), where H out is the output height, W out is the output width, the final convolutional calculation result matrix Y convout has a size of (H out , W out ); First, perform a linear transformation on the input image X to convert it into a matrix X′. After the transformation, the size of X′ is (H out ×W out , K×K×C), where each row corresponds to the expansion of a convolution window; the convolution F is converted into a matrix F′ with a size of (K×K×C, 1), and each column corresponds to the expansion of a convolution kernel; the convolution operation is represented as a two-dimensional matrix multiplication: Y = X′×F′, and the convolution algorithm is completed through two-dimensional matrix multiplication.

4. The artificial intelligence application deployment method based on domestic GPU according to claim 1, wherein In Step 4, when implementing the kernels of the pooling and activation operators using OpenCL, the steps are as follows: Allocate the optimal number of threads T for the current domestic GPU device, and the global number of threads is (H * W) / (T * 4); When it is an activation operator, the following parallel program is used for calculation, where n is the index of the number of threads: output[n] = input[n] > 0? input[n] : input[n] * 0.1 When it is a pooling operator, the following parallel program is used for calculation: output[n] = max(input[n - 3:n + 3]).

5. The artificial intelligence application deployment method based on domestic GPUs according to claim 1, wherein In Step 4, when implementing the kernel of the fully connected operator using OpenCL, the steps are as follows: Place the vector B in the local memory with faster memory access on the GPU, use OpenCL to allocate M parallel threads, perform the multiplication calculation of the array A and the vector B on the M threads, and use the reduction algorithm on the workgroup to perform parallel acceleration of the fully connected operator using shared memory.

6. The artificial intelligence application deployment method based on domestic GPU according to claim 1, wherein In Step 4, when implementing the kernel of the LSTM operator using OpenCL, the steps are as follows: (1)Concurrently compute the forget gate and the input gate, where the forget gate f t is jointly determined by the hidden layer state h t-1 at the previous time step and the input x t at the current time step: f t = σ(W f h t-1 + U f x t + b f ) Among them, W f , U f and b f are the weight and bias matrices related to the forget gate, and σ is the sigmod function; The input gate consists of the following two parts: i t = σ(W i h t-1 + U i x t + b i ) a t = σ(W a h t-1 + U a x t + b a ) Among them, W i 、W a 、U i 、U a 、b i and b a are the weight and bias matrices related to the input gate; (2) Calculate the update gate: C t = C t-1 ⊙ f t + i t ⊙ a t where ⊙ is the Hadamard product; (3) The output of the output gate is: O t = σ(W O h t-1 + U O x t + b O )。 7. The artificial intelligence application deployment method based on domestic GPUs according to claim 1, wherein In Step 4, when implementing the kernel of the convolution operator using OpenCL, the steps are as follows: Allocate the global number of threads as H * W, and use the following parallel program for calculation, where n is the index of the number of threads: output[n] = input0[n] + input1[n]; output[n] = input0[n] - input1[n]; output[n] = input0[n] * input1[n]; output[n] = input0[n] / input1[n].