A large-scale simulation and model derivation method, device, equipment and storage medium

By establishing an RPC data transmission connection between the CPU and the GPU, and setting up data compression and load balancing modules in the CPU, the problems of low GPU usage and waste of computing resources in large-scale simulation are solved, and more efficient computing resource utilization and cost reduction are achieved.

CN114153590BActive Publication Date: 2025-05-09GUANGZHOU WERIDE TECH LTD CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111214569.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-05-09
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

When performing large-scale simulation and model derivation devices, they need to use a large number of GPU machines, resulting in a huge cost of scene simulation, and the GPU usage rate is low, resulting in waste of computing resources.

Method used

By setting up simulation software modules in the CPU and establishing data transmission connections between the CPU and GPU using RPC, a large number of relatively low-cost CPUs are deployed to run simulation software, and model derivation is performed with a small number of powerful GPUs. In addition, a data compression module and a load balancing module are set up in the CPU to optimize data transmission and resource allocation.

Benefits of technology

It effectively reduces the cost of scenario simulation, increases the usage rate of GPU, saves computing resources, and realizes more efficient allocation and execution of computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114153590B_ABST
    Figure CN114153590B_ABST
Patent Text Reader

Abstract

The present application discloses a large-scale simulation and model derivation method, device, equipment and storage medium. The device includes: a CPU provided with a simulation software module and a GPU provided with a model derivation module; the number of the GPUs is less than the number of the CPUs; the CPUs are connected to the GPUs via RPC data transmission. The present application can reduce the cost of scene simulation during large-scale simulation, and is conducive to improving the utilization rate of the GPU and saving computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a large-scale simulation and model derivation method, device, equipment and storage medium. Background Art

[0002] Large-scale simulation technology based on cloud platforms plays an important role in the field of autonomous driving. It can reduce the time cost and safety hazards of algorithm testing in test vehicles, while improving the efficiency of algorithm testing and the coverage of algorithm testing for various scenarios.

[0003] Model derivation is a very important part of the autonomous driving algorithm. In scene verification, the algorithm model needs to be used as follows: 1. Deploy various algorithm models to the GPU machine; 2. Use the verification software to collect sensor data, such as various obstacle information, surrounding environment information and other data; 3. Use the verification software to obtain high-precision map data of the vehicle; 4. Use the verification software to input the collected data into the algorithm model; 5. Use the algorithm model engine to calculate and feedback the derivation results 5. Use the verification software to continue to perform simulation tasks based on the derivation results.

[0004] Autonomous driving algorithms need to run hundreds of thousands of scenarios every day to evaluate their effectiveness, generating tens of millions of model derivation requests, which places extremely high demands on model derivation.

[0005] The simulation software runs in the CPU machine, while the model derivation is performed in the GPU machine. Since the access software in the CPU machine accesses the GPU machine in the form of native code, the CPU machine and the GPU machine must be set up in a one-to-one correspondence, which results in the existing large-scale simulation and model derivation devices requiring the use of a large number of GPU machines when performing large-scale simulation, making the cost of scene access very high, and the low utilization rate of GPU, resulting in a waste of computing resources. Summary of the invention

[0006] To this end, the technical problem solved by the embodiments of the present application is to provide a large-scale simulation and model derivation device, which can reduce the cost of scene simulation during large-scale simulation, and is conducive to improving the utilization rate of the GPU and saving computing resources.

[0007] In order to solve the above technical problems, the technical solutions adopted in this application are as follows:

[0008] On the one hand, an embodiment of the present application provides a large-scale simulation and model derivation device, including: a CPU provided with a simulation software module and a GPU provided with a model derivation module; the number of the GPUs is less than the number of the CPUs; the CPU is connected to the GPU via RPC (Remote Procedure Call Protocol) data transmission.

[0009] Furthermore, the large-scale simulation and model derivation device provided in the embodiment of the present application includes: the CPU is also provided with a data compression module for running a data compression algorithm; the output end of the simulation software module is connected to the input end of the data compression module.

[0010] Preferably, the data compression algorithm is a gzip compression algorithm.

[0011] Alternatively, the data compression algorithm is the lz4 compression algorithm.

[0012] Furthermore, the CPU is also provided with a load balancing module; the output end of the simulation software module is connected to the input end of the load balancing module through the data compression module.

[0013] Preferably, the CPU is also provided with a client; the simulation software module, data compression module and load balancing module are arranged in the client; the GPU is also provided with a server and a data decompression module; the model derivation module and the data decompression module are arranged in the server; the load balancing module is used to obtain and store the address information of the server; when the client issues an RPC request, the client requests the load balancing module to obtain the address information of the server, and the client establishes a data transmission connection between the client and the server according to the address information returned by the load balancing module.

[0014] More preferably, the load balancing module is a load balancing module using an xds load balancing strategy.

[0015] Alternatively, the load balancing module is a load balancing module that uses a glb load balancing strategy.

[0016] Alternatively, the load balancing module is a load balancing module using a round_robin load balancing strategy.

[0017] On the other hand, an embodiment of the present application provides a large-scale simulation and model derivation method, including:

[0018] Build a CPU with a simulation software module;

[0019] Build GPUs with model inference modules and fewer in number than CPUs;

[0020] Use RPC to establish data transmission connection between CPU and GPU.

[0021] Furthermore, the large-scale simulation and model derivation method comprises:

[0022] Building a data compression module in the CPU for running a data compression algorithm;

[0023] Construct a data transmission connection between the output of the simulation software module and the input of the data compression module.

[0024] Furthermore, the large-scale simulation and model derivation method includes:

[0025] Build a load balancing module in the CPU;

[0026] The data compression module is used to construct a data transmission connection between the output end of the simulation software module and the input end of the load balancing module.

[0027] Preferably, the large-scale simulation and model derivation method comprises:

[0028] A client for setting a simulation software module, a data compression module, and a load balancing module is constructed in the CPU;

[0029] Build a server in the GPU to set up the model inference module and data decompression module;

[0030] When the client sends an RPC request, the client requests the load balancing module to obtain the address information of the server, and a data transmission connection is established between the client and the server according to the address information returned by the load balancing module.

[0031] On the other hand, an embodiment of the present application provides a device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the steps of any one of the above-mentioned large-scale simulation and model derivation methods when executing the computer program.

[0032] On the other hand, an embodiment of the present application provides a storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned large-scale simulation and model derivation methods are performed.

[0033] In summary, compared with the prior art, the technical solution provided in the embodiment of the present application has at least the following beneficial effects:

[0034] 1. The embodiment of the present application establishes a data transmission connection between the CPU and the GPU by utilizing the RPC method. When performing large-scale simulation, a large number of CPUs with relatively low costs compared to GPUs can be deployed to run the simulation software, and a small number of GPUs with more powerful computing power than the CPU can be used for model derivation. The existing large-scale simulation and model derivation devices can effectively reduce the cost of scene simulation, and are conducive to improving the utilization rate of GPUs and saving computing resources.

[0035] 2. The embodiment of the present application sets a data compression module in the CPU so that the data collected by the simulation software module is compressed before being input into the GPU for model derivation, thereby solving the technical problem that the network card bandwidth of the GPU where the model derivation module is located is fully occupied when model derivation is initiated concurrently in large-scale real-time access tasks.

[0036] 3. The embodiment of the present application not only establishes a long connection between the CPU and the GPU by setting a load balancing module in the CPU, but also maintains the connection status by monitoring the server status of the GPU. At the same time, the model derivation requests issued by the CPU can be sent to the server of the GPU evenly, ensuring that the loads of the server instances in the GPU are similar. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a structural schematic diagram of a large-scale simulation and model derivation device provided by the first exemplary embodiment of the present application.

[0038] Figure 2 It is a structural diagram of a large-scale simulation and model derivation device provided by the sixth exemplary embodiment of the present application.

[0039] Figure 3 It is a flowchart of a large-scale simulation and model derivation method provided by the eighth exemplary embodiment of the present application.

[0040] Figure 4 It is a partial flowchart of the large-scale simulation and model derivation method provided by the ninth exemplary embodiment of the present application.

[0041] Figure 5 It is a partial flowchart of the large-scale simulation and model derivation method provided by the tenth exemplary embodiment of the present application.

[0042] Figure 6 It is a schematic diagram of the structure of the device provided by the twelfth exemplary embodiment of the present application. DETAILED DESCRIPTION

[0043] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make modifications to the present embodiment without any creative contribution as needed, but such modifications are protected by the patent law as long as they are within the scope of the claims of the present application.

[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0045] The term "comprise" and any variations thereof in the specification and claims of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0046] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0047] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.

[0048] Figure 1 The first exemplary embodiment of the present application provides a large-scale simulation and model derivation device, which includes: a CPU provided with a simulation software module and a GPU provided with a model derivation module; the number of the GPUs is less than the number of the CPUs; the CPU is connected to the GPU via RPC data transmission.

[0049] The embodiment of the present application establishes a data transmission connection between the CPU and the GPU by utilizing the RPC method. When performing large-scale simulation, a large number of CPUs with relatively low costs compared to GPUs can be deployed to run the simulation software, and a small number of GPUs with more powerful computing power than the CPU can be used for model derivation. The existing large-scale simulation and model derivation devices can effectively reduce the cost of scene simulation, and are conducive to improving the utilization rate of GPUs and saving computing resources.

[0050] It should be noted that the number of CPUs and GPUs described in the first exemplary embodiment of the present application is not limited to Figure 1 For those skilled in the art, the specific number of CPUs and GPUs is set according to the actual number of simulation tasks, as long as the number of CPUs is greater than the number of GPUs.

[0051] A second exemplary embodiment of the present application provides a large-scale simulation and model derivation device, which Figure 1 Further improvements are made on the basis of the first exemplary embodiment shown, and the specific improvements are as follows:

[0052] The CPU is also provided with a data compression module for running a data compression algorithm; the output end of the simulation software module is connected to the input end of the data compression module.

[0053] The second exemplary embodiment of the present application sets a data compression module in the CPU so that the data collected by the simulation software module is compressed before being input into the GPU for model derivation, thereby solving the technical problem of occupancy of the network card bandwidth of the GPU where the model derivation module is located when model derivation is initiated concurrently in large-scale real-time tasks.

[0054] The third exemplary embodiment of the present application provides a large-scale simulation and model derivation device, which is further improved on the basis of the second exemplary embodiment of the present application, and the specific improvements are as follows:

[0055] The data compression algorithm is a gzip compression algorithm, thereby achieving data compression at the framework level.

[0056] At the same time, the inventors found that implementing data compression at the framework level would cause the following new technical problems when implementing the embodiments of the present application: compressing and decompressing tens of megabytes of data would introduce a delay of about 200-300ms to the GPU, and would occupy a lot of CPU resources. Therefore, in order to solve the above new technical problems, the fourth exemplary embodiment of the present application provides a large-scale simulation and model derivation device, which is further improved on the basis of the second exemplary embodiment of the present application, and the specific improvements are as follows:

[0057] The data compression algorithm is the lz4 compression algorithm, thereby achieving data compression at the application level.

[0058] The fourth exemplary embodiment of the present application uses the lz4 compression algorithm to perform data compression. Based on the observation of online data, the CPU to GPU latency is reduced by nearly 200ms, and the CPU sample data also proves that the CPU usage rate is reduced from 5% to 0.076%.

[0059] The fifth exemplary embodiment of the present application provides a large-scale simulation and model derivation device, which is further improved on the basis of the second to fourth exemplary embodiments of the present application, and the specific improvements are as follows:

[0060] The CPU is also provided with a load balancing module; the output end of the simulation software module is connected to the input end of the load balancing module through the data compression module.

[0061] The fifth exemplary embodiment of the present application not only establishes a long connection between the CPU and the GPU by setting a load balancing module in the CPU, but also can maintain the connection status by monitoring the server status of the GPU. At the same time, the model derivation requests issued by the CPU can be sent to the server of the GPU on an even basis to ensure that the loads of the server instances in the GPU are similar.

[0062] In order to describe in detail the data transmission connection between the CPU and the GPU established by using the RPC method described in the first to fifth exemplary embodiments, Figure 2 This is the sixth exemplary embodiment of the present application, which is further improved on the basis of the fifth exemplary embodiment of the present application, and the specific improvements are as follows:

[0063] The CPU is also provided with a client; the simulation software module, data compression module and load balancing module are arranged in the client; the GPU is also provided with a server and a data decompression module; the model derivation module and the data decompression module are arranged in the server; the load balancing module is used to obtain and store the address information of the server; when the client issues an RPC request, the client requests the load balancing module to obtain the address information of the server, and the client establishes a data transmission connection between the client and the server according to the address information returned by the load balancing module, thereby ensuring the connection accuracy between the client and the server.

[0064] It should be noted that: the address information includes IP information and port information; the data transmission connection is a network connection; the number of CPUs and GPUs is not limited to Figure 2 The number shown in the figure; for those skilled in the art, the specific number of CPUs and GPUs is set according to the actual number of simulation tasks, as long as the number of CPUs is greater than the number of GPUs. Each CPU contains a number of clients, and each GPU contains a number of servers. The number of clients and servers is not limited to Figure 2 The quantity displayed.

[0065] The seventh exemplary embodiment of the present application is further improved on the basis of the sixth exemplary embodiment of the present application, and the specific improvements are as follows:

[0066] The load balancing module is a load balancing module that uses the xds load balancing strategy.

[0067] The seventh exemplary embodiment of the present application can deploy the CPU and GPU on a cloud platform because its load balancing module is a load balancing module that uses an xds load balancing strategy.

[0068] In other embodiments, the load balancing module may be a load balancing module using a glb load balancing strategy.

[0069] Alternatively, in other embodiments, the load balancing module may also be a load balancing module that uses a round-robin load balancing strategy, so that all subsequent RPC requests of the client may be sent to the server through the data transmission connection in a round-robin manner.

[0070] Each module of the above-mentioned large-scale simulation and model derivation device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0071] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device described in the present application is divided into different functional units or modules to complete all or part of the functions described above.

[0072] Figure 3 The eighth exemplary embodiment of the present application provides a large-scale simulation and model derivation method, including:

[0073] Build a CPU with a simulation software module;

[0074] Build GPUs with model inference modules and fewer in number than CPUs;

[0075] Use RPC to establish data transmission connection between CPU and GPU.

[0076] The eighth exemplary embodiment of the present application establishes a data transmission connection between the CPU and the GPU by utilizing the RPC method. When performing large-scale simulation, a large number of CPUs with relatively low costs compared to GPUs can be deployed to run the simulation software, and a small number of GPUs with more powerful computing power than the CPU can be used for model derivation. The existing large-scale simulation and model derivation devices can effectively reduce the cost of scene simulation, and are conducive to improving the utilization rate of the GPU and saving computing resources.

[0077] The ninth exemplary embodiment of the present application provides a large-scale simulation and model derivation method, which Figure 3 Further improvements are made on the basis of the eighth exemplary embodiment shown in FIG. Figure 4 The specific improvements are as follows:

[0078] The large-scale simulation and model derivation method comprises:

[0079] Building a data compression module in the CPU for running a data compression algorithm;

[0080] Construct a data transmission connection between the output of the simulation software module and the input of the data compression module.

[0081] The ninth exemplary embodiment of the present application constructs a data compression module in the CPU so that the data collected by the simulation software module is compressed before being input into the GPU for model derivation, thereby solving the technical problem of occupancy of the network card bandwidth of the GPU where the model derivation module is located when model derivation is initiated concurrently in large-scale real-time tasks.

[0082] The data compression algorithm is a gzip compression algorithm, thereby achieving data compression at the framework level.

[0083] Alternatively, the data compression algorithm is the lz4 compression algorithm, thereby achieving data compression at the application level. By using the lz4 compression algorithm for data compression, based on the observation effect of online data, the CPU to GPU latency is reduced by nearly 200ms, and the CPU sample data also proves that the CPU usage rate is reduced from 5% to 0.076%.

[0084] The tenth exemplary embodiment of the present application provides a large-scale simulation and model derivation method, which Figure 4 Further improvements are made on the basis of the ninth exemplary embodiment shown in FIG. Figure 5 The specific improvements are as follows:

[0085] The large-scale simulation and model derivation method comprises:

[0086] Build a load balancing module in the CPU;

[0087] The data compression module is used to construct a data transmission connection between the output end of the simulation software module and the input end of the load balancing module.

[0088] The tenth exemplary embodiment of the present application not only establishes a long connection between the CPU and the GPU by constructing a load balancing module in the CPU, but also can maintain the connection status by monitoring the server status of the GPU. At the same time, it can also send the model derivation requests issued by the CPU to the server of the GPU on an even basis, ensuring that the loads of the server instances in the GPU are similar.

[0089] The eleventh exemplary embodiment of the present application provides a large-scale simulation and model derivation method, which Figure 5 Further improvements are made on the basis of the tenth exemplary embodiment shown, and the specific improvements are as follows:

[0090] The large-scale simulation and model derivation method comprises:

[0091] A client for setting a simulation software module, a data compression module, and a load balancing module is constructed in the CPU;

[0092] Build a server in the GPU to set up the model inference module and data decompression module;

[0093] When the client sends an RPC request, the client requests the load balancing module to obtain the address information of the server, and a data transmission connection is established between the client and the server according to the address information returned by the load balancing module.

[0094] The eleventh exemplary embodiment of the present application establishes a data transmission connection between the client and the server according to the address information returned by the load balancing module, thereby ensuring the connection accuracy between the client and the server.

[0095] It should be noted that the RPC methods described in the first to eleventh exemplary embodiments of the present application mainly include HTTP method and GRPC method, and GRPC method is more preferred, mainly based on the consideration that GRPC is a long connection method, thereby reducing the overhead of establishing a TCP connection between the client and the server each time.

[0096] Figure 6It is a device provided by the twelfth exemplary embodiment of the present application, and the device may be a server. The device includes a processor, a memory and a communication interface connected via a system bus. Among them, the processor of the device is used to provide computing and control capabilities. The memory of the device can be implemented by any type of volatile or non-volatile storage device or a combination thereof, and the volatile or non-volatile storage device includes but is not limited to: a disk, an optical disk, an EEPROM, an EPROM, a SRAM, a ROM, a magnetic storage device, a flash memory, and a PROM. The memory of the device provides an environment for the operation of the operating system and computer programs stored therein. The communication interface of the device is a network interface, and the network interface is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps of the large-scale simulation and model derivation method described in the above embodiment are implemented.

[0097] In another embodiment of the present application, a storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the large-scale simulation and model derivation method described in the above embodiment are implemented. The storage medium includes but is not limited to: ROM, RAM, CD-ROM, magnetic disk, and floppy disk.

[0098] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device described in the present application is divided into different functional units or modules to complete all or part of the functions described above.

[0099] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A large-scale simulation and model derivation device, characterized in that: include: A CPU is provided with a simulation software module and a GPU is provided with a model derivation module; the number of the GPUs is less than the number of the CPUs; the CPU is connected to the GPU via RPC data transmission; The CPU is also provided with a data compression module for running a data compression algorithm; the output end of the simulation software module is connected to the input end of the data compression module; the CPU is also provided with a load balancing module; the output end of the simulation software module is connected to the input end of the load balancing module through the data compression module; The CPU is also provided with a client; the simulation software module, data compression module and load balancing module are arranged in the client; the GPU is also provided with a server and a data decompression module; the model derivation module and the data decompression module are arranged in the server; the load balancing module is used to obtain and store the address information of the server; when the client issues an RPC request, the client requests the load balancing module to obtain the address information of the server, and the client establishes a data transmission connection between the client and the server according to the address information returned by the load balancing module.

2. The large-scale simulation and model derivation device according to claim 1, characterized in that: The data compression algorithm is the gzip compression algorithm.

3. The large-scale simulation and model derivation device according to claim 1, characterized in that: The data compression algorithm is the lz4 compression algorithm.

4. The large-scale simulation and model derivation device according to claim 1, characterized in that: The load balancing module is a load balancing module that uses the xds load balancing strategy.

5. The large-scale simulation and model derivation device according to claim 1, characterized in that: The load balancing module is a load balancing module that uses a glb load balancing strategy.

6. The large-scale simulation and model derivation device according to claim 1, characterized in that: The load balancing module is a load balancing module using a round_robin load balancing strategy.

7. A large-scale simulation and model derivation method, characterized in that: include: Build a CPU with a simulation software module; Build GPUs with model inference modules and fewer in number than CPUs; Use RPC to establish data transmission connection between CPU and GPU; Building a data compression module in the CPU for running a data compression algorithm; Establishing a data transmission connection between an output terminal of the simulation software module and an input terminal of the data compression module; Build a load balancing module in the CPU; Using the data compression module to construct a data transmission connection between the output end of the simulation software module and the input end of the load balancing module; A client for setting a simulation software module, a data compression module, and a load balancing module is constructed in the CPU; Build a server in the GPU to set up the model inference module and data decompression module; When the client sends an RPC request, the client requests the load balancing module to obtain the address information of the server, and a data transmission connection is established between the client and the server according to the address information returned by the load balancing module.

8. A device, characterized in that The method comprises a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the steps of the large-scale simulation and model derivation method as claimed in claim 7 when executing the computer program.

9. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the large-scale simulation and model derivation method as claimed in claim 7.

Citation Information

Patent Citations

  • Method, system and device for executing simulation test task and medium

    CN112148481A

  • Inference engine design method for improving GPU (Graphics Processing Unit) calculation throughput by separating script from model

    CN113342538A

  • Method and device for determining image processing mode

    CN113366531A