Simulation engine scheduling method and system for unmanned system simulation test

By designing a simulation engine scheduling method and system in multi-GPU scenarios, and automatically selecting the best-performance server using the engine scheduling components, the problem of low resource utilization in traditional deployment methods is solved, and the management of the virtualization environment is simplified, achieving efficient and convenient resource scheduling.

CN120086000APending Publication Date: 2025-06-03ZHONGKE NANJING SOFTWARE TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311612286.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In multi-GPU scenarios, traditional server deployment methods lead to low utilization of GPU simulation engine resources. At the same time, the deployment methods based on virtualization technology are complex and have high thresholds, making it difficult to efficiently and conveniently schedule simulation engine resources in multiple GPUs.

Method used

Design a simulation engine scheduling method and system for unmanned system simulation testing, automatically select the best performance server through the engine scheduling components, and schedule the simulation engine to achieve efficient and convenient resource utilization. The method includes a client sending an engine scheduling request to the engine scheduling component, obtaining GPU resource usage information of multiple servers, calculating the idle status score of each server, and selecting a target server to improve overall resource utilization.

Benefits of technology

Through real-time computing and automated selection of engine scheduling components, the utilization rate of simulation engine resources is improved, the configuration and management process is simplified, the threshold for use is lowered, and suitable for multi-GPU and virtualized environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086000A_ABST
    Figure CN120086000A_ABST
Patent Text Reader

Abstract

The invention discloses a simulation engine scheduling method and system for an unmanned system simulation test, and belongs to the technical field of simulation tests.The method comprises the steps that a client side of a simulation application program sends an engine scheduling request to an engine scheduling assembly, and the engine scheduling request comprises video memory demand quantity and a simulation test identifier; the engine scheduling component obtains GPU resource use information of a server of the simulation application program through a GPU resource monitoring agent; the engine scheduling component determines an idle state score of each server based on the video memory demand quantity and the GPU resource use information of each server; the engine scheduling component selects a target server from the servers based on the simulation test identifier and the idle state scores of all the servers; and the engine scheduling component sends an engine scheduling response to the client, wherein the engine scheduling response is used for indicating related information of the target server of the scheduled simulation engine. According to the invention, the accuracy of simulation engine scheduling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of simulation testing, and particularly to a simulation engine scheduling method and system for unmanned system simulation testing. Background Art

[0002] The powerful parallel computing ability and flexible resource management mechanism of the GPU (Graphics Processing Unit) make it an important technical tool in the field of virtual simulation technology. It can accelerate the execution of virtual simulation tasks and optimize the utilization of resources. This is of great significance for improving the efficiency and performance of virtual simulation, and also brings great potential to some key application fields (such as engineering, medicine, and weather forecasting, etc.).

[0003] In the process of the gradual development of virtual simulation application scenarios towards distribution, the scenario of multiple GPUs is becoming more and more common. When using multiple GPUs, in the traditional server deployment method, usually only one GPU is installed on one server, and only one simulation engine runs on one GPU; and if the server and GPU have high performance, multiple GPUs can be installed on one server, and virtualization technology is often used to divide a server into multiple virtual servers, and each virtual server can run one simulation engine. Virtualization technology can provide resource isolation and the independence of virtual servers, enabling multiple virtual environments to run on the same GPU simultaneously.

[0004] However, for the above traditional server deployment methods, there are often situations of resource surplus or resource idleness, resulting in low utilization rate of the simulation engine resources of the GPU; and for the above server deployment method based on virtualization technology, the complexity of configuring and managing virtual machines is relatively high, and it is necessary to understand the configuration parameters, hardware requirements, and network settings of virtual machines, etc. It is not very user-friendly for some non-professional users and raises the usage threshold. Therefore, in the scenario of multiple GPUs, there is an urgent need for a method to efficiently and conveniently schedule the simulation engine resources in multiple GPUs and improve the overall resource utilization rate. Summary of the Invention

[0005] For the unmanned system simulation testing, especially for the large-scale unmanned system simulation testing, the present invention provides a simulation engine scheduling method and system in the scenario of multiple GPUs, which can automatically select the server with the best performance through the engine scheduling component to schedule the simulation engine for the simulation test currently executed by the client, efficiently and conveniently, and improve the overall resource utilization rate.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A simulation engine scheduling method for unmanned system simulation testing, which is applied to a simulation engine scheduling system. The simulation engine scheduling system includes a client of a simulation application, a server of the simulation application, and an engine scheduling component. The method includes:

[0008] Step 1: The client of the simulation application sends an engine scheduling request to the engine scheduling component. The engine scheduling request includes the video memory demand and the simulation test identifier.

[0009] Step 2: The engine scheduling component obtains the GPU resource usage information of multiple servers of the simulation application through the GPU resource monitoring agent in the server of the simulation application.

[0010] Step 3: The engine scheduling component determines the idle state score of each server based on the video memory demand and the GPU resource usage information of each server.

[0011] Step 4: The engine scheduling component selects a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers.

[0012] Step 5: The engine scheduling component sends an engine scheduling response to the client. The engine scheduling response is used to indicate the relevant information of the target server of the scheduled simulation engine.

[0013] Optionally, the determining the idle state score of each server includes:

[0014] Step 3.1: For server S i , based on the video memory demand and the GPU resource usage information of the Kth GPU in server S i , determine whether the Kth GPU can meet the client's usage requirements for the simulation engine.

[0015] Step 3.2: If the Kth GPU cannot meet the client's usage requirements for the simulation engine, determine that the idle state score of the Kth GPU is zero, and after setting K = K + 1, start from Step 3.1 again.

[0016] Step 3.3: If the Kth GPU can meet the client's usage requirements for the simulation engine, calculate the idle state score of the Kth GPU based on the video memory demand and the GPU resource usage information of the Kth GPU.

[0017] Step 3.4: Statistically calculate the idle state scores of all GPUs in server S i to obtain the idle state score of server S i .

[0018] Optionally, the GPU resource usage information of the Kth GPU includes GPU utilization rate, video memory utilization rate, video memory usage, GPU temperature, performance status, GPU power consumption, and total GPU power;

[0019] Based on the video memory demand and the GPU resource usage information of the Kth GPU, calculate the idle status score of the Kth GPU, which is implemented through the following calculation formula:

[0020]

[0021] where W 1 and W 2 , w 3 , W 4 and W 5 are the weight values of the video memory dimension, GPU utilization rate dimension, GPU power consumption dimension, GPU temperature dimension, and performance status dimension, respectively.

[0022] Optionally, the engine scheduling component selects a target server from multiple servers based on the simulation test identifier and the idle status scores of all servers, including:

[0023] Step 4.1: Obtain the simulation test identifiers of the simulation tests currently executed by each server;

[0024] Step 4.2: Based on the simulation test identifiers corresponding to each server, determine whether there is a preferred server among all servers. The simulation test identifier corresponding to the preferred server is the same as the simulation test identifier corresponding to the client;

[0025] Step 4.3: If there is a preferred server and there is a server with a non-zero idle status score in the preferred server, select the server with the highest idle status score from the preferred servers as the target server;

[0026] Step 4.4: If there is no preferred server, or there is a preferred server but the idle status scores of all servers in the preferred server are zero, select the server with the highest idle status score from all servers as the target server.

[0027] Optionally, the engine scheduling response includes the IP address of the target server and the device identifier of the target GPU; the IP address of the target server is used for the client to send an execution request for the simulation test task to the target server.

[0028] Optionally, the engine scheduling response further includes the IP addresses and priorities of at least one backup server. The backup servers are selected from all servers with non-zero idle status scores except the target server; the higher the idle status score of the backup server, the higher the priority of the backup server.

[0029] Optionally, the method further includes:

[0030] If the target server cannot execute the simulation test task, send a task failure indication to the client;

[0031] After receiving the task failure indication, the client sends an execution request for the simulation test task to the standby server according to the IP address of the standby server; wherein, if there are multiple standby servers, according to the priority order of the multiple standby servers, first send an execution request to the standby server with a higher priority. If the standby server with a higher priority still cannot execute the simulation test task, continue to send an execution request to the standby server with a lower priority until the simulation test task is executed. The present invention also proposes a simulation engine scheduling system for unmanned system simulation testing, which is used to schedule the simulation engines running on the GPUs in the servers. The system includes: a client of the simulation application, a server of the simulation application, and an engine scheduling component;

[0032] The client and the server of the simulation application are respectively in communication connection with the engine scheduling component;

[0033] The client is used to interact with the user, receive the user's simulation test requirements, and send an engine scheduling request to the engine scheduling component. The engine scheduling request includes the video memory demand and the simulation test identifier;

[0034] The engine scheduling component is used to obtain the GPU resource usage information of multiple servers of the simulation application, determine the idle state scores of each server based on the video memory demand and the GPU resource usage information of each server, select a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers, and send an engine scheduling response to the client;

[0035] The engine scheduling component includes an information collection module and a calculation and decision module. The information collection module is responsible for establishing communication connections with the client and the server of the simulation application and performing data exchanges, and is also responsible for forwarding the GPU resource usage information and the engine scheduling request to the calculation and decision module; the calculation and decision module is responsible for determining the target server of the scheduled simulation engine based on the GPU resource usage information and the engine scheduling request to respond to the engine scheduling request of the client;

[0036] The server is used for simulation calculation and data storage. There is a GPU in the server, and a simulation engine runs on the GPU.

[0037] The beneficial effects of the present invention include:

[0038] 1. Through the engine scheduling component, the ability of the server to be scheduled by the simulation engine is calculated in real time. In the present invention, an engine scheduling component is newly added between the client and the server of the simulation application, and a GPU resource monitoring agent is installed in each server; the engine scheduling component obtains the GPU resource usage information of each server through the GPU resource monitoring agents of each GPU, and based on this, calculates in real time the ability of each server to be scheduled by the simulation engine, so as to automatically select the server with the best performance to respond to the engine scheduling request of the client, make full use of the GPU performance of each server, be efficient and convenient, and improve the overall resource utilization rate. Moreover, the engine scheduling component provided by the present invention is applicable to both physical servers and virtual servers, and the configuration and performance differences between servers do not affect the scheduling of the simulation engine, having strong applicability.

[0039] 2. Based on the simulation test identifier, preferentially select the server to be scheduled by the simulation engine that performs the same simulation test as the client. The present invention designs a simulation test identifier that uniquely identifies different simulation tests. The engine scheduling component can obtain the simulation test identifiers of the simulation tests currently executed by each server. In addition, the client encapsulates the simulation test identifier of the simulation test it is currently executing in the engine scheduling request when requesting the scheduling of the simulation engine. Thus, based on the simulation test identifier corresponding to the client and the simulation test identifiers corresponding to multiple servers, the engine scheduling component preferentially selects the server to be scheduled by the simulation engine that performs the same simulation test as the client, so as to ensure that the same simulation test is executed on the same server as much as possible, without considering data synchronization between different servers, and alleviates data latency during the simulation test process.

[0040] 3. Construct GPU resource usage information from multiple dimensions, and calculate the idle state score based on the GPU resource usage information. In the present invention, the ability of each server to be scheduled by the simulation engine is indicated by the idle state score, and the idle state score of each server is calculated by the engine scheduling component based on the GPU resource usage information of the server. The present invention constructs GPU resource usage information from multiple dimensions, not only considering GPU-related parameters for running the simulation engine such as video memory usage, performance status, GPU temperature, and GPU power, eliminating the influence of server configuration and performance differences on the scheduling of the simulation engine; but also considering the video memory demand of the client for performing the simulation test, realizing targeted calculation of whether there are sufficient remaining resources on the GPU in the server to run the simulation engine for the specific simulation test task of the client, and improving the accuracy of the simulation engine scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of a simulation engine scheduling method for unmanned system simulation testing;

[0042] Figure 2Schematic diagram of a simulation engine scheduling system for unmanned system simulation testing. Detailed implementation manners

[0043] The present invention will be further described in detail below with reference to the accompanying drawings.

[0044] In one embodiment, the present invention provides a simulation engine scheduling method for unmanned system simulation testing, as Figure 1 shown, the method includes:

[0045] Step 1: The client of the simulation application program sends an engine scheduling request to the engine scheduling component. The engine scheduling request includes the video memory demand and the simulation test identifier.

[0046] The simulation application program is used to perform unmanned system simulation testing, and it corresponds to a client and a server. Among them, the client is used for user operations and interface display, etc.; the server is used to provide background services, such as simulation calculation and data storage, etc. In the present invention, the simulation application program corresponds to multiple servers, so that it can not only support multiple simulation tests to be executed in parallel, but also support the simulation testing of large-scale unmanned systems.

[0047] When there is a simulation test requirement in any client (such as creating a new simulation test or adding a new simulation test task to the currently running simulation test), it can send an engine scheduling request to the engine scheduling component to request scheduling of the simulation engine running on the GPU in the server (such as UE (Unreal Engine), CARLA, Unity, etc.) to perform the simulation test. Optionally, the IP address of the engine scheduling component is preset in each client, and it can send an engine scheduling request to the engine scheduling component according to this IP address.

[0048] Among them, the engine scheduling request includes the video memory demand. Optionally, multiple sets of corresponding relationships between simulation test tasks and video memory demands are preset in each client. When the client receives the simulation test task requested by the user, it obtains the corresponding video memory demand based on the above corresponding relationship and encapsulates the video memory demand in the engine scheduling request. By pre-establishing the corresponding relationship between the simulation test task and the video memory demand, the user does not need to have the ability to evaluate and calculate the corresponding video memory demand, and the simulation test task set by the user can be automatically converted into the video memory demand, reducing the usage threshold. It should be understood that the simulation testing described in the present invention includes one or more simulation test tasks, and new simulation test tasks can be added at any time according to requirements during the operation of the simulation test. Among them, the simulation test refers to the test project, and the simulation test task refers to the specific simulation content in the test project. For example, the simulation test can be an unmanned aerial vehicle cluster simulation test project, and the unmanned aerial vehicle cluster simulation test project can include simulation test tasks such as 20 fixed-wing unmanned aerial vehicles and 30 quadrotor unmanned aerial vehicles.

[0049] In the present invention, the engine scheduling request further includes a simulation test identifier. The simulation test identifier is used to uniquely identify different simulation tests, and it can be an ID sequence or the like. Optionally, project management of all simulation tests within the simulation application can be performed on one server, and all clients of the simulation application are connected to this server. When a user initiates a new request for a simulation test through a certain client, this client obtains the simulation test identifier of this simulation test from the server, and encapsulates this simulation test identifier in the engine scheduling request.

[0050] Step 2: The engine scheduling component obtains the GPU resource usage information of multiple servers of the simulation application through the GPU resource monitoring agent.

[0051] The engine scheduling component includes an information collection module, and each server includes a GPU resource monitoring agent. The information collection module establishes connections with the GPU resource monitoring agents of multiple servers respectively to obtain the GPU resource usage information of the servers through the GPU resource monitoring agents. Among them, the GPU resource monitoring agent can be implemented as a GPU monitoring tool (such as nvidia-smi, etc.) or an API (such as Nvidia's NvAPI or AMD's ADL, etc.).

[0052] In the present invention, one or more GPUs can be installed in one server, and each GPU corresponds to a GPU resource usage information. Thus, the GPU resource usage information of each server includes the GPU resource usage information of each GPU in this server. Among them, the GPU resource usage information of each GPU includes but is not limited to: GPU utilization rate, video memory utilization rate, video memory usage amount, GPU temperature, performance status, GPU power usage amount, and total GPU power.

[0053] Optionally, the GPU resource monitoring agent of each server can collect the GPU resource usage information of this server at a preset time interval (such as 0.1 second, etc.), and send the GPU resource usage information to the information collection module of the engine scheduling component in real time, or feedback the GPU resource usage information to it based on the request of the information collection module. In addition, before sending the GPU resource usage information to the information collection module, the GPU resource monitoring agent can also perform processing such as formatting on the initially obtained GPU resource usage information, so that the GPU resource usage information can be recognized and processed by the engine scheduling component. Among them, the engine scheduling component also includes a calculation and decision-making module. For the parsing of the GPU resource usage information, it can be directly parsed after being obtained by the information collection module and the parsed GPU resource usage information is sent to the calculation and decision-making module for subsequent processing; or it can be immediately forwarded to the calculation and decision-making module after being obtained by the information collection module, and the calculation and decision-making module parses it and then performs subsequent processing.

[0054] Step 3: Based on the video memory demand and the GPU resource usage information of each server, the engine scheduling component determines the idle state score of each server.

[0055] In the present invention, the calculation decision module in the engine scheduling component determines the idle state score of each server based on the video memory demand of the client and the GPU resource usage information of each server. The present invention does not limit the calculation method of the idle state score. Optionally, the number of GPUs with video memory remaining greater than or equal to the video memory demand in the server can be determined first based on the video memory demand and the GPU resource usage information of each server, and then the idle state score of the server can be determined based on the number of GPUs; or, the idle state score of each GPU in the server can be determined first based on the video memory demand and the GPU resource usage information of each server, and then the idle state score of the server can be determined based on the idle state scores of all GPUs in the server; or, the number of GPUs with video memory remaining less than the video memory demand in the server can be determined first based on the video memory demand and the GPU resource usage information of each server, and the idle state scores of each GPU with video memory remaining greater than or equal to the video memory demand in the server can be determined, and then the idle state score of the server can be determined based on the number of GPUs and the idle state scores of the corresponding GPUs.

[0056] In one example, taking the idle state score of each server being calculated based on the idle state scores of all GPUs in the server as an example, the above Step 3 includes the following sub-steps (Steps 3.1 to 3.4).

[0057] Step 3.1: For server S i , based on the video memory demand and the GPU resource usage information of the Kth GPU in server S i , determine whether the Kth GPU can meet the client's usage requirements for the simulation engine.

[0058] Server S i is the ith server among multiple servers of the simulation application program, and i is a positive integer. For any server of the simulation application program, the calculation methods of Steps 3.1 to 3.4 can be adopted to calculate the idle state score of the server. Optionally, in the present invention, the multiple servers of the simulation application program can be sorted according to the number of GPUs in the server or the total video memory of the GPUs in the server, etc., or based on a random policy, or in combination with the sending time or receiving time of the GPU resource usage information of the server, and then the idle state scores of each server are calculated in sequence.

[0059] For server S iFor each GPU in the server, the idle state score of the GPU is calculated based on the video memory demand and the GPU resource usage information of the GPU. i If multiple GPUs are installed in the server, the server S can be allocated based on the total size of the GPU memory or the GPU memory usage, or based on a random strategy. i The multiple GPUs are sorted, and then the idle state score of each GPU is calculated in sequence.

[0060] For server S i In the Kth GPU, K is a positive integer. When calculating the idle state score of the Kth GPU, it is necessary to first determine whether the Kth GPU can meet the client's demand for the use of the simulation engine based on the client's video memory demand and the resource usage information of the Kth GPU. If the Kth GPU cannot meet the client's demand for the use of the simulation engine, execute the following step 3.2, otherwise execute the following step 3.3.

[0061] The present invention does not limit the specific method of determining whether the GPU can meet the usage requirements. Optionally, one or more dimensions of operating conditions can be set, and according to the degree of influence of these operating conditions on the GPU operating effect, whether the GPU meets the current operating conditions is judged in descending order of the degree of influence. If the GPU meets the current operating conditions, continue to judge whether the GPU meets the next operating condition, until it is judged that the GPU meets the operating conditions of all dimensions, then it is determined that the GPU meets the usage requirements; if the GPU does not meet the current operating conditions, then it is determined that the GPU does not meet the client's usage requirements for the simulation engine.

[0062] Exemplarily, taking the GPU resource usage information of the Kth GPU including GPU usage rate, video memory usage rate, video memory usage, GPU temperature, performance status, GPU power usage, and GPU total power as an example, the operating conditions of dimensions such as video memory, GPU usage rate, GPU power, GPU temperature, and performance status can be set, and according to the degree of influence of these operating conditions on the GPU operating effect, in the order of influence from high to low (such as the order of remaining video memory, video memory usage rate, performance status, GPU temperature, and GPU power consumption rate), it is determined whether the Kth GPU meets the current operating conditions.

[0063] That is, first determine whether the remaining video memory of the Kth GPU is greater than or equal to the video memory demand. If the remaining video memory is less than the video memory demand (the current operating conditions are not met), it is determined that the Kth GPU does not meet the client's usage requirements for the simulation engine. If the remaining video memory is greater than or equal to the video memory demand (the current operating conditions are met), then continue to determine whether the video memory utilization rate of the Kth GPU is less than or equal to the video memory utilization rate threshold; if the video memory utilization rate is greater than the video memory utilization rate threshold (the current operating conditions are not met), it is determined that the Kth GPU does not meet the client's usage requirements for the simulation engine. If the video memory utilization rate is less than or equal to the video memory utilization rate threshold (the current operating conditions are met), then continue to determine whether the performance state of the Kth GPU is less than or equal to the performance state threshold; if the performance state is greater than the performance state threshold (the current operating conditions are not met), it is determined that the Kth GPU does not meet the client's usage requirements for the simulation engine. If the performance state is less than or equal to the performance state threshold (the current operating conditions are met), then continue to determine whether the GPU temperature of the Kth GPU is less than or equal to the GPU temperature threshold; if the GPU temperature is greater than the GPU temperature threshold (the current operating conditions are not met), it is determined that the Kth GPU does not meet the client's usage requirements for the simulation engine. If the GPU temperature is less than or equal to the GPU temperature threshold (the current operating conditions are met), then continue to determine whether the GPU power consumption rate of the Kth GPU is less than or equal to the power consumption rate threshold; if the GPU power consumption rate is greater than the power consumption rate threshold (the current operating conditions are not met), it is determined that the Kth GPU does not meet the client's usage requirements for the simulation engine. If the GPU power consumption rate is less than or equal to the power consumption rate threshold (the current operating conditions are met), it is determined that the Kth GPU meets the client's usage requirements for the simulation engine.

[0064] Before the above determination, since the GPU resource usage information of the Kth GPU includes the video memory utilization rate, the video memory usage, the GPU power consumption rate, and the total GPU power, after parsing the GPU resource usage information of the Kth GPU, the calculation formulas for the remaining video memory and the GPU power consumption rate of the Kth GPU are as follows.

[0065]

[0066]

[0067] Of course, in practical applications, the GPU resource usage information of each GPU may also directly include the remaining video memory and the GPU power consumption rate of the GPU, so that the remaining video memory and the GPU power consumption rate can be directly obtained by parsing the GPU resource usage information of the GPU.

[0068] Step 3.2: If the Kth GPU cannot meet the client's demand for the simulation engine, determine that the idle state score of the Kth GPU is zero, and after setting K = K + 1, start over from the above Step 3.1.

[0069] If the Kth GPU cannot meet the current client's demand for the simulation engine, the computing decision module determines that under the current engine scheduling request of the current client, the idle state score of the Kth GPU is zero. After that, the computing decision module can set K = K + 1 and start over from the above Step 3.1 to continue determining whether the next GPU can meet the current client's demand for the simulation engine.

[0070] Among them, if server S i includes M GPUs, where M is a positive integer. After setting K = K + 1, the computing decision module compares K with M. If K is less than or equal to M, it means that the GPUs in server S i have not been traversed completely, and it can start over from the above Step 3.1; if K is greater than M, it means that the GPUs in server S i have been traversed completely, and continue to execute the following Step 3.4.

[0071] Step 3.3: If the Kth GPU can meet the client's demand for the simulation engine, calculate the idle state score of the Kth GPU based on the video memory demand and the GPU resource usage information of the Kth GPU.

[0072] If the Kth GPU can meet the current client's demand for the simulation engine, in order to further compare to obtain the optimal GPU and server for simulation engine scheduling, the computing decision module further calculates the idle state score of the Kth GPU based on the video memory demand and the GPU resource usage information of the Kth GPU. The present invention does not limit the calculation method of the idle state score. Optionally, according to the specific content of the GPU resource usage information, the resource margin of the GPU running the simulation engine can be calculated from multiple dimensions, and the resource margins of multiple dimensions can be statistically processed to obtain the idle state score of the GPU.

[0073] Exemplarily, the GPU resource usage information of the Kth GPU includes GPU utilization rate, video memory utilization rate, video memory usage, GPU temperature, performance state, GPU power consumption, and total GPU power; thus, the resource margins of dimensions such as video memory, GPU power, GPU temperature, and performance state can be calculated, and after normalizing the resource margins of these dimensions respectively and then weighted summing, the idle state score of the Kth GPU can be determined. For example, the calculation formula for the idle state score of the Kth GPU is as follows. Among them, W 1 、W 2 、W3 , W 4 and W 5 are the weight values for the dimensions of video memory, GPU utilization rate, GPU power consumption, GPU temperature, and performance status respectively. Their sum is equal to 1. In practical applications, the weight values for each dimension can be flexibly set in combination with the influence degree of each dimension of resources on the simulation test running effect, the user's emphasis on each dimension of resources, etc.

[0074]

[0075] Step 3.4: Statistically process the idle state scores of all GPUs in server S i to obtain the idle state score of server S i .

[0076] For server S i , repeat the above steps 3.1 to 3.3 until the idle state scores of all GPUs in server S i are calculated. Then, statistically process the idle state scores of all GPUs to calculate the idle state score of server S i . The present invention does not limit the specific manner of the above statistical processing. Optionally, the calculation and decision-making module can directly calculate the average of the idle state scores of all GPUs; or, the idle state scores of all GPUs can be weighted and summed, and the weight values for this weighted summation process can be determined based on factors such as the total video memory of the GPU and the remaining video memory; or, first count the number of GPUs with non-zero idle state scores, and then use the proportion of this number of GPUs in the total number of GPUs in server S i as the weight and multiply it by the sum of the idle state scores of all GPUs.

[0077] Exemplarily, taking the statistical processing process that combines the number of GPUs and the idle state scores of GPUs as an example, the calculation formula for the idle state score of server S i is as follows.

[0078]

[0079]

[0080] Step 4: The engine scheduling component selects a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers.

[0081] Among them, the target server is scheduled by the simulation engine to execute simulation tests and the simulation test tasks therein. In the present invention, on the one hand, based on the idle state scores of all servers, the server with the highest idle state score is selected as the target server to ensure that the best-performing server is selected for the simulation test currently executed by the client to schedule the simulation engine; on the other hand, based on the simulation test identifier of the simulation test, the server that executes the same simulation test as the client is preferentially selected as the target server, as much as possible to ensure that the same simulation test is executed on the same server, alleviating data latency during the simulation test process.

[0082] In one example, step 4 above includes the following sub-steps (steps 4.1 to 4.4).

[0083] Step 4.1: Obtain the simulation test identifiers of the simulation tests currently executed by each server.

[0084] The calculation decision module in the engine scheduling component obtains the simulation test identifiers corresponding to each server. Optionally, when it is necessary to obtain the simulation test identifier, the calculation decision module can obtain the latest simulation test identifier corresponding to each server in real time through the information collection module in the engine scheduling component; or, the calculation decision module maintains an identifier mapping table, which includes at least the correspondence between one server and the simulation test identifier, and the simulation test identifier corresponding to each server can be directly obtained through this identifier mapping table.

[0085] For the case where the calculation decision module maintains an identifier mapping table, whenever the calculation decision module determines that a certain server is scheduled by the simulation engine for a certain simulation test, a correspondence is established between the server and the simulation test identifier of the simulation test, so as to update the identifier mapping table. In addition, whenever a certain simulation test ends, the client that newly creates the simulation test can forward the end instruction of the simulation test to the calculation decision module through the information collection module. When the calculation decision module receives the end instruction, the correspondence between all the servers that execute the simulation test and the simulation test identifier of the simulation test is deleted from the identifier mapping table to avoid the identifier mapping table being too long and facilitate subsequent rapid determination of the preferred server.

[0086] Step 4.2: Based on the simulation test identifiers corresponding to each server, select the preferred server with the same simulation test identifier as the client from multiple servers.

[0087] In the present invention, if a user creates a certain simulation test through a client and a need for a new simulation test task arises during the running of this simulation test, the server that executes this simulation test is preferentially selected to schedule a simulation engine for this simulation test task. Thus, after the calculation and decision-making module obtains the simulation test identifier corresponding to the client by parsing the engine scheduling requirement of the client, it compares the simulation test identifier corresponding to the client with the simulation test identifiers corresponding to each server. If the simulation test identifier corresponding to a certain server is the same as the simulation test identifier corresponding to the client, it means that this server and the client execute the same simulation test, and this server is used as the preferred server. It should be understood that the preferred server may include one or more servers. For example, for the simulation test of a large-scale unmanned system, a simulation test may require multiple servers to execute simultaneously. Through the above method, it is possible to preferentially execute the same simulation test on the same server and alleviate data delay during the simulation test process.

[0088] Of course, when a user creates a certain simulation test through a client and initiates the simulation engine scheduling requirement for the simulation test task for the first time, there is temporarily no server among multiple servers that executes this simulation test, that is, there is no preferred server corresponding to this simulation test; after scheduling a simulation engine for the first simulation test task of this simulation test, when scheduling a simulation engine for other simulation test tasks of this simulation test subsequently, there is a preferred server corresponding to this simulation test among multiple servers. Therefore, in the case where a preferred server can be selected through step 4.2, the following step 4.3 or 4.4 is continued; otherwise, the following step 4.4 is continued.

[0089] Step 4.3: If there is a preferred server and there is a server among the preferred servers with a non-zero idle state score, select the server with the highest idle state score from the preferred servers as the target server.

[0090] In the case where there is a preferred server, the calculation and decision-making module further determines whether the idle state scores of all the preferred servers are zero. If there is a preferred server with a non-zero idle state score, the preferred server with the highest idle state score is used as the target server to implement the scheduling of the simulation engine through this target server; otherwise, step 4.4 is executed as follows. Among them, in the case where there is only one preferred server and the idle state score of this preferred server is non-zero, this preferred server is directly used as the target server.

[0091] Step 4.4: If there is no preferred server, or there is a preferred server but the idle state scores of all the servers among the preferred servers are zero, select the server with the highest idle state score from all the servers as the target server.

[0092] In the absence of a preferred server, the computing decision module selects the server with the highest idle state score from multiple servers as the target server to ensure load balancing among multiple servers as much as possible. In addition, in the presence of a preferred server, if the idle state scores of all preferred servers are zero, it means that among the servers performing the same simulation test as the client, all GPUs do not have sufficient remaining resources to run the simulation engine to execute the simulation test tasks generated by the client. Therefore, to ensure the successful scheduling of the simulation engine for this simulation test task, the computing decision module selects the server with the highest idle state score from multiple servers as the target server.

[0093] It should be understood that if multiple servers adopt the traditional deployment method, after determining that the target server is scheduled with the simulation engine, the target server distributes the simulation test tasks to specific GPUs according to its pre-set algorithm; if virtualization technology is adopted during the deployment of multiple servers, after determining that the target server is scheduled with the simulation engine, the computing decision module can further determine which GPU in the target server will execute the simulation test task. The present invention does not limit the specific determination method of the GPU. Optionally, the computing decision module can count the idle state scores of all GPUs in each server to obtain the idle state score of the server. Thus, after determining the target server according to the idle state scores of multiple servers, based on the idle state scores of each GPU in the target server, the GPU with the highest idle state score is selected as the target GPU, and subsequently, the simulation engine is scheduled for the simulation test tasks generated by the client with the target GPU in the target server.

[0094] Step 5: The engine scheduling component sends an engine scheduling response to the client, and the engine scheduling response is used to indicate the relevant information of the target server of the scheduled simulation engine.

[0095] The engine scheduling response is used to indicate the relevant information of the target server of the scheduled simulation engine. Optionally, the engine scheduling response includes the IP address of the target server, so that the client can send an execution request for the simulation test task to the target server according to this IP address subsequently. In addition, for the case where virtualization technology is adopted during the deployment of multiple servers, if the computing decision module also determines the target GPU in the target server, the engine scheduling response includes not only the IP address of the target server but also the device identifier of the target GPU, such as the ID sequence, etc.

[0096] It should be noted that in practical applications, after the client sends an execution request for a simulation test task to the target server, the target server may be unable to effectively execute the simulation test task due to reasons such as machine failures or insufficient server memory. To address this situation, the engine scheduling response sent by the engine scheduling component to the client may further include the IP addresses of at least one backup server. The backup server can be all the servers with a non-zero idle state score except the target server among the preferred servers in step 4.3 above or among the multiple servers in step 4.4 above; in the case of selecting multiple backup servers, the engine scheduling response may further include the priorities of these multiple backup servers. The higher the idle state score of the backup server, the higher its priority.

[0097] Based on this, if the target server is unable to effectively execute the simulation test task, it can send a task failure indication to the client. After receiving the task failure indication, the client can send an execution request for the simulation test task to the backup server according to the IP addresses of the backup servers included in the engine scheduling response. Among them, if there are multiple backup servers, the execution request is sent to the backup server with a higher priority in the order of the priorities of the multiple backup servers. If the backup server with a higher priority is still unable to effectively execute the simulation test task, the execution request is continued to be sent to the backup server with a lower priority until the simulation test task is effectively executed. Of course, in practical applications, both the target server and the backup server in the engine scheduling response may be unable to effectively execute the simulation test task, then the client can modify the engine scheduling request and send the modified engine scheduling request to the engine scheduling component to re-execute the scheduling process.

[0098] Optionally, after the above step 3, the calculation and decision module of the engine scheduling component can further determine whether the idle state scores of multiple servers are all zero. If there is a server with a non-zero idle state score among the multiple servers, the above step 4 is continued; if the idle state scores of all multiple servers are zero, there is no server with remaining GPU resources to run the simulation engine to execute the simulation test task generated by the client. The engine scheduling component can return a scheduling failure indication to the client. After receiving the scheduling failure indication, the client can modify the engine scheduling request and send the modified engine scheduling request to the engine scheduling component, or can send the original engine scheduling request to the engine scheduling component again after a period of time.

[0099] In another embodiment, the present invention proposes a simulation engine scheduling system for unmanned system simulation testing corresponding to the method of the above embodiment, as Figure 2 shown. The simulation engine scheduling system includes a client of the simulation application, a server of the simulation application, and an engine scheduling component;

[0100] The client and server of the simulation application respectively establish communication connections with the engine scheduling component;

[0101] The client is used to interact with the user, receive the user's simulation test requirements, and send an engine scheduling request to the engine scheduling component. The engine scheduling request includes the video memory demand and the simulation test identifier;

[0102] The engine scheduling component is used to obtain the GPU resource usage information of multiple servers of the simulation application, determine the idle state scores of each server based on the video memory demand and the GPU resource usage information of each server, select a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers, and send an engine scheduling response to the client;

[0103] The engine scheduling component includes an information collection module and a calculation and decision-making module. The information collection module is responsible for establishing communication connections with the client and server of the simulation application and performing data exchanges, and is also responsible for forwarding the GPU resource usage information and the engine scheduling request to the calculation and decision-making module; The calculation and decision-making module is responsible for determining the target server of the scheduled simulation engine based on the GPU resource usage information and the engine scheduling request to respond to the engine scheduling request of the client;

[0104] The server is used for simulation calculation and data storage. A GPU is provided in the server, and the simulation engine runs on the GPU.

[0105] For the introduction and description of other functions, working methods, etc. of each component in the simulation engine scheduling system, please refer to the introduction and description in the above embodiments of the simulation engine scheduling method, and will not be elaborated here.

[0106] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A simulation engine scheduling method for unmanned system simulation testing, characterized in that, it is applied to a simulation engine scheduling system, and the simulation engine scheduling system includes a client of a simulation application, a server of the simulation application, and an engine scheduling component; the method includes: Step 1, the client of the simulation application sends an engine scheduling request to the engine scheduling component, and the engine scheduling request includes the video memory demand and the simulation test identifier; Step 2, the engine scheduling component obtains the GPU resource usage information of multiple servers of the simulation application through the GPU resource monitoring agent in the server of the simulation application; Step 3, the engine scheduling component determines the idle state score of each server based on the video memory demand and the GPU resource usage information of each server; Step 4, the engine scheduling component selects a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers; Step 5, the engine scheduling component sends an engine scheduling response to the client, and the engine scheduling response is used to indicate the relevant information of the target server of the scheduled simulation engine.

2. The simulation engine scheduling method for unmanned system simulation testing according to claim 1, characterized in that, the determining the idle state score of each server includes: Step 3.1: For server S i , based on the video memory demand and the GPU resource usage information of the K-th GPU in server S i , determine whether the K-th GPU can meet the client's usage requirements for the simulation engine; Step 3.2: If the Kth GPU cannot meet the client's usage requirements for the simulation engine, determine that the idle state score of the Kth GPU is zero, and after setting K = K + 1, start from Step 3.1 and execute again; Step 3.3: If the Kth GPU can meet the client's usage requirements for the simulation engine, calculate the idle state score of the Kth GPU based on the video memory demand and the GPU resource usage information of the Kth GPU; Step 3.4: Statistic server S i Obtain the idle state scores of all GPUs in the server S i to get the idle state score of the server S 3. The simulation engine scheduling method for unmanned system simulation testing according to claim 2, characterized in that, the GPU resource usage information of the Kth GPU includes GPU utilization rate, video memory utilization rate, video memory usage, GPU temperature, performance status, GPU power usage, and total GPU power; the calculating the idle state score of the Kth GPU based on the video memory demand and the GPU resource usage information of the Kth GPU is realized through the following calculation formula: Among them, w 1 , w 2 , w 3 , w 4 and w 5 are the weight values of the video memory dimension, GPU utilization rate dimension, GPU power consumption dimension, GPU temperature dimension, and performance state dimension, respectively.

4. The simulation engine scheduling method for unmanned system simulation testing according to claim 1, characterized in that, the engine scheduling component selects a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers, including: Step 4.1: Obtain the simulation test identifiers of the simulation tests currently executed by each server; Step 4.2: Based on the simulation test identifiers corresponding to each server, determine whether there is a preferred server among all servers, and the simulation test identifier corresponding to the preferred server is the same as the simulation test identifier corresponding to the client; Step 4.3: If there is a preferred server and there is a server with a non-zero idle state score in the preferred server, select the server with the highest idle state score from the preferred servers as the target server; Step 4.4: If there is no preferred server, or there is a preferred server but the idle state scores of all servers in the preferred server are zero, then select the server with the highest idle state score from all servers as the target server.

5. The simulation engine scheduling method for unmanned system simulation testing according to claim 1, wherein, the engine scheduling response includes the IP address of the target server and the device identifier of the target GPU; the IP address of the target server is used for the client to send an execution request for the simulation test task to the target server.

6. The simulation engine scheduling method for unmanned system simulation testing according to claim 5, wherein, the engine scheduling response further includes the IP addresses and priorities of at least one backup server, and the backup server is selected from all servers with non-zero idle state scores except the target server; the higher the idle state score of the backup server, the higher the priority of the backup server.

7. The simulation engine scheduling method for unmanned system simulation testing according to claim 6, wherein, the method further includes: if the target server fails to execute the simulation test task, send a task failure indication to the client; after receiving the task failure indication, the client sends an execution request for the simulation test task to the backup server according to the IP address of the backup server; wherein, if there are multiple backup servers, according to the priority order of the multiple backup servers, first send the execution request to the backup server with a higher priority, and if the backup server with a higher priority still fails to execute the simulation test task, continue to send the execution request to the backup server with a lower priority until the simulation test task is executed.

8. A simulation engine scheduling system for unmanned system simulation testing, wherein, the simulation engine scheduling system includes: a client of the simulation application, a server of the simulation application, and an engine scheduling component; the client and the server of the simulation application are respectively in communication connection with the engine scheduling component; the client is used to interact with the user, receive the user's simulation test requirements, and send an engine scheduling request to the engine scheduling component, and the engine scheduling request includes the video memory demand and the simulation test identifier; the engine scheduling component is used to obtain the GPU resource usage information of multiple servers of the simulation application, determine the idle state scores of each server based on the video memory demand and the GPU resource usage information of each server, select a target server from multiple servers based on the simulation test identifier and the idle state scores of all servers, and send an engine scheduling response to the client; the engine scheduling component includes an information collection module and a calculation and decision module. The information collection module is responsible for establishing communication connections with the client and the server of the simulation application and performing data exchanges, and is also responsible for forwarding the GPU resource usage information and the engine scheduling request to the calculation and decision module; the calculation and decision module is responsible for determining the target server of the scheduled simulation engine based on the GPU resource usage information and the engine scheduling request to respond to the engine scheduling request of the client. The server is used for simulation calculation and data storage. A GPU is provided in the server, and a simulation engine runs on the GPU.