A system that improves GPU computing efficiency in large data volumes and high concurrency scenarios

By deploying haproxy and thrift services between the client and the CPU server, concurrent requests for GPU graphics cards are amplified, and the problem of low computing efficiency of GPU graphics cards in large data volume and high concurrency scenarios is solved, efficient data processing is achieved and resource waste is avoided.

CN114237922BActive Publication Date: 2025-05-13SHIQU INTERACTIVE (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111096450.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-17
Publication Date
2025-05-13
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

In the case of large data volume and high concurrency, it is difficult for the existing technology to effectively improve the computing efficiency of GPU graphics cards, resulting in low utilization of GPU graphics cards, wasted computing power, and the risk of memory overflow and calculation timeout.

Method used

By deploying open source haproxy request forwarding service and thrift RPC service between the client and the CPU server, concurrent request amplification of GPU graphics cards is realized, and request polling and forwarding is used to ensure that the GPU graphics card remains in full load operation state.

Benefits of technology

It realizes the computing efficiency of GPU graphics cards without adding additional high-performance hardware resources, avoiding the risk of memory overflow and computing timeout, and improving data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114237922B_ABST
    Figure CN114237922B_ABST
Patent Text Reader

Abstract

The present invention discloses a system for improving the computing efficiency of a GPU graphics card in a large data volume and high concurrency scenario, comprising: a client, a CPU server, three servers for RPC services and a GPU graphics card, wherein the client is used to send concurrent data requests, the CPU server is used to receive data processing requests and forward the request polling to a subsequent processing end, and the GPU graphics card has a model computing service. The present application does not add additional high-performance hardware resources, utilizes a relatively cheap CPU server, and the client only needs to send a relatively small amount of request data each time, so that the GPU graphics card can reach a full-load computing state, thereby improving data processing efficiency, and is not prone to overflow and timeout risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of GPU graphics card computing efficiency, and in particular to a system for improving the computing efficiency of a GPU graphics card in large data volumes and high concurrency scenarios. Background Art

[0002] After the deep learning model training is completed, it is generally necessary to deploy an inference service on the GPU graphics card to provide calculation results based on the deep learning model for data requests sent by the client. For example, after training the text public opinion classification model, it is necessary to deploy an inference service on the GPU graphics card to quickly provide calculation results for text public opinion classification requests sent by the client.

[0003] In order to improve the utilization of GPU graphics cards in scenarios with large amounts of data and high concurrent requests, so as to increase the data processing speed of inference services, the commonly used strategies in the industry are:

[0004] 1. The client increases the number of data sent in each request (batch size), or increases the request concurrency through multi-threading and multi-processing to put all the computing pressure on the GPU graphics card. However, due to performance and bandwidth limitations, most clients using CPU servers generally find it difficult to significantly increase the request concurrency, which will cause the GPU graphics card utilization rate to not reach saturation, wasting the GPU's computing power and not easily improving the overall data processing efficiency; and if the amount of data sent in each request is increased, the GPU graphics card will need to construct a larger data matrix to accommodate and process this data, which may increase the risk of video memory overflow and the risk of a single request calculation timeout.

[0005] 2. Add more SSDs or deploy more GPU graphics cards, that is, to increase the speed of data reading, writing and computing by horizontally expanding expensive high-performance hardware resources. But this will obviously increase the hardware cost of the system, which is not a friendly and realistic solution for the majority of start-up or R&D teams with relatively limited funds. Moreover, if the computing power of a single GPU graphics card cannot be fully utilized, horizontally expanding the number of GPU graphics cards will also cause greater waste of resources. Summary of the invention

[0006] The purpose of the present invention is to solve the above problems and to propose a system for improving the computing efficiency of GPU graphics cards in large data volumes and high concurrency scenarios.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A system for improving the computing efficiency of a GPU graphics card in large data volumes and high concurrency scenarios, comprising: a client, a CPU server, three servers for RPC services and a GPU graphics card, wherein the client is used to send concurrent data requests, the CPU server is used to receive data processing requests and forward the request polling to subsequent processing ends, and the GPU graphics card has a model computing service.

[0009] Preferably, the CPU server deploys an open source haproxy request forwarding service to receive requests from clients and forward them to the RPC service in a polling manner.

[0010] Preferably, the three servers for RPC services all use open source thrift services as RPC services, further amplifying the concurrent request degree to the GPU graphics card.

[0011] Preferably, the CPU server receives the computing request with amplified concurrency sent by the RPC service started by thrift, and forwards it to the model computing service started on the GPU graphics card by polling.

[0012] Preferably, the GPU graphics card starts the computing services of three public opinion models to actually calculate the public opinion classification task of the text.

[0013] Preferably, after the model calculation service is completed, the result can be returned to the client in the reverse order of the process number. The open source deployment framework adopted above supports two-way data request and transmission.

[0014] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0015] This application does not add additional high-performance hardware resources, utilizes relatively cheap CPU servers, and each client only needs to send a relatively small amount of request data each time, and supports more client connections, so that the GPU graphics card can reach the full load computing state, thereby improving data processing efficiency and not prone to risks such as overflow and timeout. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of the system flow structure provided according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] See also Figure 1 , the present invention provides a technical solution:

[0019] A system for improving the computing efficiency of GPU graphics cards in large data volumes and high concurrency scenarios includes: a client, a CPU server, three servers for RPC services, and a GPU graphics card. The client is used to send concurrent data requests, the CPU server is used to receive data processing requests and forward the requests to subsequent processing ends in a polling manner, and the GPU graphics card has a model computing service. The CPU server deploys an open source haproxy request forwarding service to receive requests from the client and forward them to the RPC service in a polling manner. The three servers for RPC services all use an open source thrift service as an RPC service to further amplify the concurrent request degree to the GPU graphics card. The CPU server receives the amplified concurrent computing request sent by the RPC service started with thrift, and forwards it to the model computing service started on the GPU graphics card in a polling manner. The GPU graphics card starts the computing services of three public opinion models to actually calculate the public opinion classification task of the text. After the model computing service completes the calculation, the result can be returned to the client in the reverse order of the process number. The above-mentioned open source deployment frameworks all support two-way data requests and transmissions.

[0020] Specifically, Figure 1 As shown in the figure, the CPU server deploys the open source haproxy request forwarding service to receive requests from clients and forward the RPC service in subsequent 3 by polling. The request forwarding configuration of the example is as follows:

[0021] listen sentiment_thrift_60000

[0022] bind 192.168.12.2:60000

[0023] mode tcp

[0024] Balance Round Robin

[0025] option tcplog

[0026] option abortonclose

[0027] option forwardfor except 127.0.0.0 / 8

[0028] maxconn 51200

[0029] server nlp0l 192.168.12.11:60001check inter 2000rise 2fall 3

[0030] server nlp02 192.168.12.11:60002check inter 2000rise 2fall 3 ......

[0031] server nlp09 192.168.12.11:60009 check inter 2000 rise 2 fall 3

[0032] server nlp010 192.168.12.12:60001 check inter 2000 rise 2 fall 3

[0033] server nlp011 192.168.12.12:60002 check inter 2000 rise 2 fall 3 ......

[0034] server nlp018 192.168.12.12:60009 check inter 2000 rise 2 fall 3

[0035] server nlp019 192.168.12.13:60001 check inter 2000 rise 2 fall 3

[0036] server nlp020 192.168.12.13:60002 check inter 2000 rise 2 fall 3 ......

[0037] server nlp027 192.168.12.13:60009check inter 2000rise 2fall 3

[0038] The IP address of the haproxy forwarding server is 192.168.12.2, and port 60000 is open to receive data requests from clients. Then, in round robin mode, the requests are forwarded to the three servers in the configuration (IP addresses are 192.168.12.11, 192.168.12.12, and 192.168.12.13). In addition, nine ports from 60001 to 60009 are opened on these three servers to receive and process requests, for a total of 3*9=27 RPC working ports.

[0039] Specifically, Figure 1As shown in the figure, the three servers used for RPC services all use the open source thrift service as the RPC service, further amplifying the concurrent requests to the GPU graphics card. The three servers used for RPC services are the three servers that receive the haproxy forwarding requests mentioned above (the IPs are 192.168.12.11, 192.168.12.12, and 192.168.12.13, and the ports from 60001 to 60009 are opened respectively). On these three servers, the open source thrift service can be used as the RPC service to connect to the client request in 1. (forwarded by haproxy polling in 2.) and the subsequent GPU graphics card inference service in 5. (also forwarded by haproxy polling in 4.). The reason why the thrift service for RPC needs to be deployed here is that the concurrent requests to the GPU graphics card can be further amplified through the multi-process (forking) or thread pool (thread pool) service mode of the thrift service. That is to say, on the one hand, it can accommodate more concurrent client requests, and on the other hand, it can request computing services on the GPU graphics card with a further enlarged concurrency relative to the client. And when the amount of data calculated by each client request is not large, by making full use of the time slice and computing performance of the GPU graphics card calculation, while keeping the GPU graphics card in full load operation, the data processing speed can be greatly improved. The 27 thrift service ports on the three RPC servers here can each support tens to hundreds of response concurrencies through settings, so as to achieve 27*N request concurrency sent to subsequent processing services. Here, N is the response concurrency configured by the program of each thrift service port (which can be set to tens to hundreds). In addition, the data processing strategy can also be configured here to divide the large amount of data sent by the client at one time into multiple batches, and the amount of data in each batch is relatively small, so as to avoid the risk of overflow and timeout caused by the subsequent GPU graphics card service due to the large amount of data processed at a single time; at the same time, through a larger concurrency, the GPU computing performance and time slice are fully utilized to keep it in a full load computing state, so as to achieve the effect of improving the data calculation speed as a whole.

[0040] Specifically, Figure 1 As shown in the figure, the CPU server receives the calculation request sent by the RPC service started by thrift, and forwards it to the model calculation service started on the GPU graphics card through polling. The GPU graphics card starts the calculation service of three public opinion models to actually calculate the public opinion classification task of the text. Here you can use the same haproxy server as in 2. The forwarding configuration of the example is as follows:

[0041] listen sentiment_gpu_service_51000

[0042] bind 192.168.12.2:51000

[0043] mode tcp

[0044] Balance Round Robin

[0045] option tcplog

[0046] option abortonc Lost

[0047] option forwardfor except 127.0.0.0 / 8

[0048] maxconn 51200

[0049] server nlp0l 192.168.12.69:50001check inter 2000rise 2fall 3

[0050] server nlp02 192.168.12.69:50002check inter 2000rise 2fall 3

[0051] server nlp03 192.168.12.69:50003check inter 2000rise 2fall 3

[0052] That is, open port 51000 on the haproxy forwarding server (IP is 192.168.12.2) to receive 27*N requests from the 27 RPC ports in 3., and then forward them to the three model inference service ports started on the GPU graphics card through round robin mode (here the GPU IP is 192.168.12.69, and the three model service ports are 50001, 50002, and 50003). Start the computing services of three public opinion models on one GPU graphics card (here Bert's three-layer model is used, and the semantic vector dimension is 768) to actually calculate the public opinion classification task of the text. The IP and three service ports of the GPU server here are 192.168.12.69 and 50001, 50002, and 50003 configured in 4 above.

[0053] Specifically, Figure 1 As shown, after the model calculation service completes the calculation, the result can be returned to the client in the reverse order of the process number. The open source deployment frameworks used above all support two-way data request and transmission.

[0054] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A system that improves GPU computing efficiency in large data volumes and high concurrency scenarios, including: A client, a CPU server, three servers for RPC services and a GPU graphics card, wherein the client is used to send concurrent data requests, the CPU server is used to receive data processing requests and forward the requests to subsequent processing ends by polling, the GPU graphics card has a model computing service, the CPU server deploys an open source haproxy request forwarding service to receive requests from the client and forward them to the RPC service by polling, the three servers for RPC services all use the open source thrift service as the RPC service, further amplifying the concurrent request degree for the GPU graphics card, the CPU server receives the computing request sent by the RPC service started by thrift, and forwards it to the model computing service started on the GPU graphics card by polling.

2. The system for improving the computing efficiency of GPU graphics cards in large data volumes and high concurrency scenarios according to claim 1, characterized in that: The GPU graphics card starts the computing services of three public opinion models to actually calculate the public opinion classification task of the text.

3. The system for improving the computing efficiency of GPU graphics cards in large data volumes and high concurrency scenarios according to claim 2, characterized in that: After the model calculation service completes the calculation, the result is returned to the client in the reverse order of the process number. The open source deployment frameworks adopted above all support two-way data request and transmission.

Citation Information

Patent Citations

  • Artificial intelligence AI system and data processing method

    CN111427702A