Simulation testing methods, equipment, storage media and products
By parsing task information and building process dependencies, and combining real-time monitoring of process status for dynamic scheduling, the problem of low testing efficiency and reliability of distributed simulation systems is solved, realizing an automated and orderly testing process, and improving testing efficiency and result accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING OPTOKO MICROELECTRONICS TECH CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, distributed simulation systems have low testing efficiency and poor reliability, mainly due to the difficulty of collaborative testing among multiple server nodes and the lack of a unified process collaboration management mechanism, resulting in inaccurate test results and low efficiency.
By parsing the task information of the test tasks, the topology of the matching distributed simulation system is determined, the dependencies between processes are constructed, and dynamic scheduling is performed in conjunction with real-time monitoring of process status to ensure that the test tasks are executed automatically and in an orderly manner, avoiding incorrect startup order or resource conflicts.
It significantly improves the testing efficiency of distributed simulation systems, enhances the accuracy and reliability of test results, reduces reliance on manual intervention, and achieves efficient and stable automated testing.
Smart Images

Figure CN122489393A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of automated testing technology, and in particular relates to a simulation testing method, equipment, storage medium and product. Background Technology
[0002] With the development of distributed technology, distributed simulation systems are widely used in various data processing and simulation scenarios. Therefore, comprehensive and efficient testing of distributed simulation systems has become a key aspect of ensuring their operational performance.
[0003] Typically, testing of distributed simulation systems is conducted using a combination of manual and semi-automated methods. However, testing the collaborative operation of multiple server nodes involved in distributed simulation systems is particularly difficult, resulting in low testing efficiency and low reliability of test results.
[0004] Therefore, how to improve the testing efficiency and reliability of distributed simulation systems is an urgent problem to be solved in this field. Summary of the Invention
[0005] This application provides a simulation testing method, device, storage medium, and product that can improve the testing efficiency and reliability of distributed simulation systems and has wide applicability.
[0006] A first aspect of this application provides a simulation testing method, comprising: acquiring task information of a test task; determining the topology of a target simulation system adapted to the task type corresponding to the task information based on the preset topology of different simulation systems corresponding to different task types; constructing a dependency relationship between processes for executing the test task based on the topology of the target simulation system, wherein the processes are processes in the server of the target simulation system; and scheduling processes to execute the test task based on the dependency relationship and the process state of the processes.
[0007] A second aspect of this application provides a testing apparatus for a simulation system, comprising: an acquisition module for acquiring task information of a test task; a determination module for determining the topology of a target simulation system adapted to the task type corresponding to the task information, based on preset topologies of different simulation systems corresponding to different task types; a construction module for constructing dependencies between processes for executing the test task based on the topology of the target simulation system, wherein the processes are processes in the server of the target simulation system; and a scheduling module for scheduling processes to execute the test task based on the dependencies and the process states of the processes.
[0008] A third aspect of the embodiments of this application provides an electronic device, the device comprising: a memory and a program or instructions stored in the memory and executable on a processor, wherein when the program or instructions are executed by the processor, they implement the simulation testing method provided in any aspect of the embodiments of this application described above.
[0009] A fourth aspect of the embodiments of this application provides a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, they implement the simulation testing method provided by any aspect of the embodiments of this application described above.
[0010] A fifth aspect of the embodiments of this application provides a computer program product, wherein when the instructions in the computer program product are executed by the processor of an electronic device, the electronic device performs the simulation testing method provided in any aspect of the embodiments of this application described above.
[0011] The simulation testing method provided in this application analyzes the task information of the test tasks to determine the topology of the matching distributed simulation system, thereby clarifying the deployment architecture and resource distribution of the simulation system. Based on this, dependencies between processes are constructed according to the topology to ensure that the tests follow the logical order within the system, avoiding test failures due to incorrect startup order or resource conflicts. Furthermore, by monitoring the process status in real time and dynamically scheduling processes based on dependencies, the test tasks can be automated and orderly advanced, reducing reliance on manual intervention and significantly improving testing efficiency. Simultaneously, through adaptive scheduling based on the system architecture, the consistency between the testing process and the actual operating environment of the system under test is ensured, enhancing the accuracy and reliability of the test results. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a simulation testing method provided in one embodiment of this application is shown; Figure 2 This illustration shows a flowchart of the automatic generation process from a JSON configuration file to a YAML configuration file according to an embodiment of this application; Figure 3 A schematic diagram of a scenario combination matrix for parameterized testing provided in one embodiment of this application is shown; Figure 4 A schematic diagram of process state control provided in one embodiment of this application is shown; Figure 5A flowchart illustrating a simulation testing method provided in one embodiment of this application is shown; Figure 6 This illustration shows a flowchart of a fault recovery process implemented through different fault recovery strategies according to an embodiment of this application; Figure 7 A schematic diagram of the network protocol and communication layer architecture provided in one embodiment of this application is shown; Figure 8 A schematic diagram of a bandwidth calculation process provided in one embodiment of this application is shown; Figure 9 This illustration shows a schematic diagram of a data processing flow provided in one embodiment of this application; Figure 10 A flowchart illustrating the determination of a performance optimization strategy according to an embodiment of this application is shown; Figure 11 A schematic diagram of the architecture of an automated testing platform for a distributed simulation system provided in one embodiment of this application is shown; Figure 12 This paper illustrates a communication network topology diagram of a distributed simulation system provided in one embodiment of the present application. Figure 13 This illustration shows a flowchart of the dynamic generation of a target configuration file according to an embodiment of this application; Figure 14 This illustration shows a schematic diagram of the content composition of a JSON configuration file provided in one embodiment of this application; Figure 15 This illustration shows a schematic diagram of the content composition of a YAML configuration file provided in one embodiment of this application; Figure 16 This illustration shows a schematic diagram of the content composition of a dynamic configuration script provided in one embodiment of this application; Figure 17 A timing diagram illustrating concurrent execution of a target configuration file provided in one embodiment of this application is shown; Figure 18 A flowchart illustrating the testing of a simulation system according to an embodiment of this application is shown; Figure 19 This is a schematic diagram of the structure of a simulation testing device provided in one embodiment of this application; Figure 20 This is a schematic diagram of an electronic device for a simulation testing apparatus provided in one embodiment of this application. Detailed Implementation
[0014] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0015] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0016] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0017] With the development of distributed technologies, distributed simulation systems are widely used in various data processing and simulation scenarios. For example, distributed simulation systems can be applied to complex scenarios in the semiconductor field, such as simulating the physical processes of chip manufacturing and the hardware-software co-verification of large-scale integrated circuits. Therefore, comprehensive and efficient testing of distributed simulation systems and / or testing tasks has become a key aspect of ensuring system performance.
[0018] Typically, testing of distributed simulation systems and their tasks is conducted using a combination of manual and semi-automated methods. Because distributed simulation systems involve the collaborative operation of multiple server nodes, manual coordination of these servers is necessary. Specifically, distributed simulation systems rely on a local control server and multiple remote servers to operate collaboratively. During testing, the orderly start-up, shutdown, and status management of various processes on each remote server are required. However, this process usually relies on manual intervention, lacking a unified process coordination management mechanism. Improper handling of process dependencies can easily lead to startup failures, and there is no automated fault detection and recovery capability when process anomalies occur. This results in low testing efficiency and poor reliability of test results for distributed simulation systems. Furthermore, test tasks usually need to be manually converted into a format readable by the distributed simulation system. However, manual conversion can introduce formatting errors, missing parameters, or low efficiency, further complicating the reliability of the test results.
[0019] In view of this, this application provides a simulation testing method, device, storage medium, and product. By parsing the task information of the test task, the topology of the matching distributed simulation system is determined, thereby clarifying the deployment architecture and resource distribution of the system under test. Based on this, dependencies between processes are constructed according to the topology to ensure that test execution follows the logical order within the system, avoiding test failures due to incorrect startup order or resource conflicts. Furthermore, by monitoring process status in real time and dynamically scheduling processes based on dependencies, test tasks can be automated and orderly advanced, thereby reducing reliance on manual intervention and significantly improving testing efficiency. Specific embodiments of the simulation testing method, device, storage medium, and product provided in this application are described below. First, the simulation testing method is introduced.
[0020] Figure 1 A flowchart illustrating a simulation testing method provided in one embodiment of this application is shown. Figure 1 As shown, the method is applied to the test center server and includes steps S110 to S140.
[0021] S110, obtain the task information of the test task.
[0022] S120, based on the preset topology of different simulation systems corresponding to different task types, determines the topology of the target simulation system that matches the task information for the corresponding task type.
[0023] S130, based on the topology of the target simulation system, constructs the dependencies between processes used to execute test tasks, where the processes are processes in the server of the target simulation system.
[0024] S140 schedules processes to execute test tasks based on dependencies and process states.
[0025] exist Figure 1 In the illustrated embodiment, by parsing the task information of the test task, the topology of the matching distributed simulation system is determined, thereby clarifying the deployment architecture and resource distribution of the system under test. Based on this, dependencies between processes are constructed according to this topology to ensure that test execution follows the logical order within the system, avoiding test failures due to incorrect startup order or resource conflicts. Furthermore, by monitoring process status in real time and dynamically scheduling processes based on dependencies, test tasks can be automated and orderly advanced, reducing reliance on manual intervention and significantly improving test efficiency. Simultaneously, adaptive scheduling based on the system architecture ensures consistency between the test process and the actual operating environment of the system under test, enhancing the accuracy and reliability of test results. It is understood that the embodiments of this application achieve efficient, stable, and repeatable automated testing of distributed simulation systems, effectively supporting the continuous integration and quality assurance of distributed simulation systems.
[0026] For example, the simulation testing method can be applied to a test center server. The test center server can be a control node in a distributed simulation system responsible for the unified management and scheduling of the entire testing process. For instance, the test center server can be a local server within the distributed simulation system.
[0027] In some embodiments, in step S110, task information for the test task can be obtained. A test task can represent one or a set of tests that the distributed simulation system needs to perform. For example, a test task may include functional verification, performance testing, and stability testing of a node in the distributed simulation system, or the entire distributed simulation system.
[0028] For example, task information can be used to describe the specific content of a test task and serves as the basis for the distributed simulation system to execute the test task.
[0029] In one example, the test center server can obtain one or more test tasks, and for each test task, obtain the corresponding task information to perform the test.
[0030] In some optional embodiments, in step S110, to quickly and accurately obtain the task information of the test tasks, an initial configuration file of the test tasks can be received and parsed to obtain a test task list containing at least one test task. The test tasks contain initial test task information stored in a preset data format. Further, for each test task in the test task list, the initial test task information is transformed according to a target data structure to obtain a target configuration file corresponding to the test task. The task information of the test task is obtained by reading the target configuration file.
[0031] The initial configuration file contains task information for the test tasks to be executed, stored in a user-readable format. The initial configuration file can be uploaded by the user to the test center server, or generated by the test center server based on user-input requirements.
[0032] In one example, the initial configuration file could be a JavaScript Object Notation (JSON) configuration file.
[0033] By parsing the initial configuration file, a list of test tasks operable by the computer program is obtained. This list includes one or more test tasks, and the initial task information for each task can be stored in a preset data format. For example, the initial configuration file can be parsed using Python libraries to obtain the test task list, and the initial task information for each test task in the list conforms to the data structure specifications of the corresponding Python libraries.
[0034] The target data structure can be a data format that each node in the distributed simulation system can directly recognize. For each test task, the initial test task information is transformed according to the target data format to obtain task information that can be directly recognized by each node.
[0035] In one example, the target data structure can be a YAML (Yin't a Markup Language) structure. YAML is a data serialization format that supports comments and combines human readability with machine parsing efficiency. There is a one-to-one correspondence between test tasks and target configuration files; that is, each test task has a corresponding YAML-formatted target configuration file.
[0036] In one example, the test center server can deploy a configuration generation module based on a configuration generation engine. The parser in the configuration generation engine reads an initial configuration file in JSON structure and parses it to obtain a list of test tasks. This list of test tasks can be represented as [job_1, job_2, ..., job_n], where job_i represents the i-th test task, i∈[1,n].
[0037] For each test task, the data format can be converted using the converter in the generation engine. For example, the initial test task information corresponding to the i-th test task contains server information. The converter can determine the host information (including IP address and port number) based on the server information, then extract the camera parameters from the current initial test task information, and finally integrate all the information to construct a complete YAML structure of task information.
[0038] Furthermore, by configuring the file generator in the generation engine, for each test task, a target configuration file (i.e., a YAML configuration file) is generated based on the task information of the corresponding YAML structure, resulting in n target configuration files. These generated target configuration files can be stored in a specified directory according to a preset naming format. For example, the preset naming format could be config_{job_name}.yml.
[0039] In some embodiments, the initial configuration file may contain a root object. The root object is the top-level data structure of the initial configuration file, used to encapsulate all relevant global configurations and task definitions in a single test task. It typically exists as a single JSON object and is the sole root node accessed by the parser. For example, the root object named "dfi_simulator" represents the configuration container for the entire test task, containing global configuration parameters for all test tasks in a single test. For example, the global configuration parameters include a "timeout_seconds" field, used to set the timeout for the test task. For instance, setting it to 3600 seconds means that each test task cannot execute for more than one hour, and will automatically terminate if the timeout occurs. Furthermore, the root object uses an array field named "hosts" to define the server topology participating in the test. Each element in this array is a server object containing two attributes: "ip" and "port," specifying the server's IP address and communication port, respectively. For example, only one local server can be configured with an IP address of 127.0.0.1 and a port of 50001. Understandably, in practical applications, multiple server objects can be added to the array according to the scale of the distributed simulation system under test to simulate a multi-machine collaborative deployment environment.
[0040] In other embodiments, the initial configuration file may also contain task objects. Task objects are objects under the root object used to organize specific test tasks, aggregating configurations for different test types or datasets in key-value pairs, with each key corresponding to an independent set of test tasks. For example, a task object named "jobs" represents a container for all test tasks, organizing multiple test tasks in key-value pairs. Each key represents the name of a test task, such as "pattern-12bit", "singleswath_12bit_800_pattern_8swath", etc., which identify the dataset characteristics or test type used by the test task; the value corresponding to each key is an array containing multiple camera configuration objects. Each camera object describes specific parameters of a simulation data source, including: the camera's unique identifier "id", the scan index "swath" (e.g., -1 indicates the default value), the image directory level "image_dir_level", the dataset storage path "data_dir" (e.g., " / mnt / share / data / ..."), and the first scan with scan direction identifier "is_first_swath_scan_dir_forward" (1 indicates forward). In this way, a test task can define multiple camera data sources to simulate multi-input scenarios.
[0041] Using the initial configuration file described above as an example, the detailed conversion process from the initial configuration file to the target configuration file will be illustrated as follows: First, after reading the JSON configuration file, the parser extracts the server topology information and the complete task list. Then, the converter iterates through each test task in the test task list, generating a separate YAML configuration file for each test task. The generated YAML configuration file can use "dfi_simulator" as the root object, which retains the "timeout_seconds" and "hosts" fields inherited from JSON to ensure that each test task is aware of its runtime environment; simultaneously, the camera array corresponding to the current test task is directly mapped to the "cameras" field in the YAML. In this way, each YAML configuration file contains the server topology and the camera parameters corresponding to that task, forming a YAML configuration file that can be directly read and executed by the simulation program.
[0042] In this embodiment, multiple test tasks are automatically extracted by receiving and parsing configuration files with a preset data format. For each test task, a format conversion is performed to generate a corresponding target configuration file. Finally, the task information for each test task is obtained by reading these target configuration files. It is understood that the above method achieves automated generation of test configurations, improving efficiency while reducing manual intervention. This avoids formatting errors or parameter omissions that may be introduced by manually writing configuration files, thus improving test reliability.
[0043] In some alternative embodiments, in step S110, technicians can manually configure the target configuration file according to the test requirements, so that the test center server can obtain the task information corresponding to the test task based on the target configuration file.
[0044] In some embodiments, in step S120, the topology of the target simulation system that matches the task information for the corresponding task type can be determined based on the preset topology of different simulation systems corresponding to different task types.
[0045] For example, a task type matching the task information can be determined from multiple preset task types, and the topology corresponding to that task type can be used as the topology of the target simulation system. The target simulation system can be a simulation system used to run test tasks.
[0046] For example, there is a correspondence between task types and the topology of the simulation system, and this correspondence can be preset by technicians. The topology can be used to describe the physical or logical connections between servers in the simulation system. Understandably, the topology determines how test tasks are distributed across one or more servers.
[0047] In one example, as shown in Table (1) below, the technicians define four test types, including: tests based on different datasets to be tested (A), GPU-based tests (B), CPU thread-based distributed tests (C), and tests based on runtime models (D). The topologies corresponding to the different test types are as follows: Table (1) Specifically, in Table (1) above, Type A tests are comparative tests based on different datasets. Type A tests aim to evaluate the performance stability of the simulation system when processing diverse test datasets. In the tests, the server topology is fixed with two servers, the number of threads is kept at 40, and GPU acceleration is not enabled. By changing the single variable of the dataset, the processing efficiency of the distributed simulation system under different data volumes or different data characteristics can be compared, thereby determining the performance of the distributed simulation system.
[0048] Category B tests are GPU-accelerated performance tests. To accurately measure the processor's improvement in simulation speed, the test environment is simplified to a single server and uses a single dataset. In this server topology, the number of GPUs is the core test variable and can be adjusted between one and two. By comparing the test results in pure CPU mode (0 GPUs) with different numbers of GPUs, the benefits of GPU acceleration and the scalability of multi-GPU parallel processing can be quantitatively analyzed.
[0049] Type C testing is a distributed scalability test based on CPU threads. Type C testing aims to evaluate the horizontal scalability of a distributed simulation system in multi-core processor and multi-node environments. The server topology itself becomes a test variable, switching between one or two servers; simultaneously, the number of CPU threads on each server is dynamically adjusted from 4 to 64. Type C testing comprehensively evaluates the processing performance of a distributed simulation system in single-machine multi-core and multi-machine collaborative scenarios, revealing the impact of resource contention, thread scheduling, and network communication on overall simulation efficiency.
[0050] Type D tests are comparative tests based on operating modes. To verify the impact of different operating modes on system performance, Type D tests use the same hardware configuration as Type A tests: two servers, 40 threads, and no GPU. The core difference lies in switching the operating mode from "non-continuous scan" to "continuous scan." By comparing the results with those of the preceding tests, the impact of this software logic-level change in scanning mode on the throughput and response latency of the distributed simulation system can be accurately assessed.
[0051] It is understandable that the four types of test tasks shown in Table (1) reflect the correspondence between the server topology and the test objectives of the distributed simulation system under test by using the dataset, number of threads, number of GPUs and running mode as configurable variables.
[0052] For type C tests, different numbers of threads and servers can be automatically combined to form various test scenarios. Furthermore, a corresponding target configuration file can be generated for each test scenario. Figure 2 This illustration shows a flowchart of the automatic generation process from a JSON configuration file to a YAML configuration file according to an embodiment of this application; Figure 3 This diagram illustrates a scenario combination matrix for parametric testing provided in one embodiment of this application. For example... Figure 2As shown, in S210, the JSON configuration file is read and parsed to obtain test task information, which includes the number of threads and the number of servers. In S220, multiple test scenarios are obtained by automatically combining the number of threads and the number of servers. Threads and servers can each be considered as a dimension to construct a test matrix to represent various test scenarios. Figure 3 As shown, the thread count dimension 310 can be configured with 8, 16, 32, and 64 threads; the server count dimension 320 can be configured with 1 or 2 servers. By combining the elements in the thread count dimension 310 and the server count dimension 320, multiple test scenarios 330 are formed, including 8×1, 8×2, 16×1, 16×2, 32×1, 32×2, 64×1, and 64×2. Among them, 8×1, 8×2, and 16×1 belong to low-load test scenarios; 16×2 and 32×1 belong to medium-load test scenarios; 32×2 and 64×1 belong to high-load test scenarios; and 64×2 belongs to extremely high-load test scenarios. Furthermore, in S230, a corresponding YAML configuration file is generated for each test scenario. The generated YAML configuration files for different test scenarios can be named according to the corresponding test scenario. Understandably, by separating the management of server topology from specific task content (such as test scenarios), testers can flexibly adjust the combination of test scenarios (such as the arrangement of different numbers of threads and servers), quickly build large-scale test matrices, and lay a solid foundation for subsequent automated execution and performance comparison analysis.
[0053] For example, when technicians build the configuration file, they can specify the server (including IP address and port number) corresponding to the test task. Alternatively, the test center server can automatically assign a server to execute the test task based on the test task.
[0054] In some embodiments, in step S130, the dependencies between processes for performing test tasks are constructed based on the topology of the target simulation system.
[0055] Here, "process" refers to the process within the server of the target simulation system.
[0056] For example, a process can be a program instance running on various servers of the target simulation system. Examples include an algorithm service (algo_service), a scheduling service (schedule_service), and a simulation simulator (Simulator). It is understood that each process is responsible for performing a specific function and works in conjunction with other processes.
[0057] For example, dependencies between processes can characterize the mutual dependence between processes at startup, operation, or shutdown. For instance, an application service can only function properly after its dependent underlying services have started. Understandably, dependencies determine the execution order of processes.
[0058] In some embodiments, in step S140, the process is scheduled to execute the test task based on the dependencies and the process state.
[0059] For example, process status can be used to describe the current running status of a process, such as "running," "stopped," or "abnormally terminated." The test center server monitors process status in real time through a local process manager or a remote process manager, which serves as the basis for scheduling processes to execute test tasks.
[0060] In one example, process states can include initial state (INT), starting state (STARTING), running state (RUNNING), error detection state (ERROR_DETECT), restarting state (RESTARTING), stopping state (STOPPING), and stopped state (STOPPED). Furthermore, process state switching can be controlled via events or commands, thereby enabling process control and state determination. Figure 4 This illustration shows a schematic diagram of process state control provided in one embodiment of this application, such as... Figure 4 As shown, process state transitions can be triggered by events, and, as Figure 4The triggering conditions are marked with arrows: When a process is in the INIT state (410), it enters the STARTING state (420) upon receiving a "start command". When the process is in the STARTING state (420) and "startup successful", it enters the RUNNING state (430); or when the process is in the STARTING state (420) and "startup failed", it enters the FAILED state (440). When the process is in the RUNNING state (430), if an "abnormal event" is detected, it enters the ERROR_DETECT state (450) and restarts according to the "automatic restart policy", while the process enters the RESTARTING state (460). After the process executes the "restart", it returns to the STARTING state (420). When the process is in the RUNNING state (430), if a "stop command" is received, it enters the STOPPING state (470), and further, when a "termination command" is received, it enters the STOPPED state (480). Furthermore, when the process is in the RUNNING state (430), during normal process operation, performance monitoring and health checks can be performed, and logs can be recorded. When a process is in the ERROR_DETECT state 450, exception detection can be performed, such as checking for process errors, network errors, and resource errors. Therefore, by sending different events or control commands to the process, the process state can be controlled and its state can be detected.
[0061] For example, in a process execution order determined based on dependencies, the current process is scheduled to execute a test task if it is determined that the previous process of the current process is in a normal state.
[0062] exist Figure 1 In the illustrated embodiment, by parsing the task information of the test task, the topology of the matching distributed simulation system is determined, thereby clarifying the deployment architecture and resource distribution of the system under test. Based on this, dependencies between processes are constructed according to this topology to ensure that test execution follows the logical order within the system, avoiding test failures due to incorrect startup order or resource conflicts. Furthermore, by monitoring process status in real time and dynamically scheduling processes based on dependencies, test tasks can be automated and orderly advanced, reducing reliance on manual intervention and significantly improving test efficiency. Simultaneously, adaptive scheduling based on the system architecture ensures consistency between the test process and the actual operating environment of the system under test, enhancing the accuracy and reliability of test results. It is understood that the embodiments of this application achieve efficient, stable, and repeatable automated testing of distributed simulation systems, effectively supporting the continuous integration and quality assurance of distributed simulation systems.
[0063] To achieve collaborative management of various processes, as another implementation method of this application, this application also provides another implementation method of the simulation testing method, as detailed in the following embodiments.
[0064] Figure 5 A flowchart illustrating a simulation testing method provided in one embodiment of this application is shown. Figure 5 As shown, the method includes steps S141 to S144: S141, Based on the topology of the target simulation system, the test task is split to obtain at least one test subtask corresponding to the server of the target simulation system.
[0065] For example, a test subtask can be the test task that has been broken down into test tasks on each server according to the server topology. It is understood that breaking down the test task into server-specific test subtasks according to the topology ensures that each server only executes the test tasks relevant to it.
[0066] In one example, the test center server may include a process management module, and the process management module may include a master controller. Test tasks can be decomposed using the task decomposition module in the master controller to obtain test subtasks corresponding to each server.
[0067] S142, Based on the dependencies and test subtasks, generate process control instructions for each of the multiple processes that execute the test subtasks.
[0068] For example, process control instructions can be commands targeting multiple processes used to execute test subtasks. For instance, the startup of process B in a server depends on process A. If process B needs to execute test subtasks, then the processes used to execute the test subtasks include both process A and process B. Accordingly, process control instructions can include control instructions for process A and control instructions for process B. For example, process control instructions for executing test subtasks can include control instructions to start process A, control instructions to start process B, and control instructions to cause process B to execute the test subtasks.
[0069] Process control instructions can be commands used to control the state of a process and / or control the execution of test subtasks by a process. For example, process control instructions can be state adjustment instructions such as controlling the start, stop, and restart of a process, or they can be instructions to control the receiving of data and start the algorithm program to execute data processing tasks.
[0070] S143, determine the priority order of multiple process control instructions based on dependencies.
[0071] For example, the priority order of process control instructions can be used to characterize the execution order of process control instructions. The priority order of process control instructions can be obtained by sorting the process control instructions according to the dependencies between processes.
[0072] In one example, during the sorting of process control instructions, the instructions can be sorted in the order of basic services, application services, and business services to obtain the priority order of multiple process control instructions.
[0073] In another example, the priority order of each process control instruction can be determined by the execution timing planning module in the main controller.
[0074] In some optional embodiments, the distributed simulation system may include a local server (i.e., a test center server) and at least one remote server. The processes for executing test tasks include: a main process deployed on the test center server to provide basic services and a simulation process for generating simulation data; a control process deployed on at least one remote server for communication and scheduling control, and an algorithm process for processing the simulation data. Multiple process control instructions include: a first process control instruction for starting the main process, a second process control instruction for starting the simulation process, a third process control instruction for starting the control process, and a fourth process control instruction for starting the algorithm process. Based on dependencies, the priority order of the multiple process control instructions, from highest to lowest, is determined as: the first process control instruction, the second process control instruction, the third process control instruction, and the fourth process control instruction.
[0075] In one example, the main controller may also include a Local Process Manager and a Remote Process Manager. A distributed simulation system may include a local server and two remote servers. The local server includes a main process (DIFService) for providing basic services and a simulation process (Simulator) for generating simulation data. Each remote server includes a control process for communication and scheduling control (Management Control Network, MCN) and an algorithm process for processing the simulation data.
[0076] The local process manager can first send a first process control command to start DIFService. Then, after determining that DIFService is running, the local process manager sends a second process control command to start the Simulator.
[0077] The remote process manager can send control commands to the first and second remote servers respectively to start the third process of the MCN in the first remote server and the third process of the MCN in the second remote server. Furthermore, the remote process manager sends a fourth process control command to the MCN in each remote server, enabling the MCN to control the process to receive images from the Simulator and start the algorithm program to process the images.
[0078] The third process control command for sending the startup MCN to the first remote server can be constructed as follows. This third process control command includes environment build control commands, startup control commands, and status feedback control commands. Specifically, the remote process manager first sends an environment build control command to the MCN to enable the MCN to build its environment. This command sets the environment variables required for MCN operation. For example, the MCN can set PATH and LD_LIBRARY_PATH according to the environment build control command, ensuring that the MCN can find dependent libraries and load the MCN executable and its dependent libraries based on those libraries, avoiding startup failures due to path issues. Further, the remote process manager sends a startup control command to control the MCN startup: connecting to the first remote server through a preset communication channel and switching to the MCN's working directory using the startup control command. This allows the relative paths used in subsequent commands to be resolved based on the working directory, ensuring that the MCN can correctly load configuration files, access dependent resources, and control the execution of the startup process control command. Furthermore, the remote process manager verifies process status through status feedback quality: After the MCN executes the start control command, the first remote server determines whether it has an MCN process and records the corresponding process identifier (PID). The first remote server can also return the check results through a preset communication channel, reporting the MCN status (e.g., running, abnormal, etc.) to the remote process manager, along with the corresponding PID for that MCN process.
[0079] It is understandable that, similar to the process described above of sending the process control instruction to the first remote server to start the MCN, other process control instructions in this application embodiment can also be constructed according to the corresponding process requirements, and this application will not elaborate on this further.
[0080] For example, process control instructions can be sent serially according to their priority order. The same priority level can include one or more process control instructions, and these instructions can be sent serially or in parallel.
[0081] In this embodiment, the main process corresponding to the first process control instruction serves as a basic service, providing runtime environment support for the entire simulation system. It is given the highest priority to ensure that the infrastructure is ready before test startup. The simulation process corresponding to the second process control instruction relies on the environment provided by the main process to generate simulation data, and is therefore placed with the next highest priority, starting immediately after the basic service is ready. The control process corresponding to the third process control instruction and the algorithm process corresponding to the fourth process control instruction are deployed on remote servers, responsible for communication scheduling and data processing. They rely on the data stream generated by the simulation process and are therefore given relatively lower priorities. This priority order determined by process function and dependencies fundamentally avoids problems such as process crashes, connection failures, or resource contention caused by disordered startup order, ensuring the stable startup and reliable operation of the distributed simulation system in collaborative testing.
[0082] S144: Execute process control instructions according to the priority order of multiple process control instructions and the process states of multiple processes.
[0083] For example, the main controller executes process control instructions sequentially according to a defined priority order, and monitors the process status in real time during execution. Instructions of the next priority are only executed if the instruction of the previous priority is executed successfully and the corresponding process status is normal.
[0084] In one example, the local process manager, upon determining that DIFService is in a "running" state, sends a process control command to the Simulator to start the Simulator. Further, the remote process manager sends process control commands to the MCNs on each remote server, enabling the MCNs to control the processes receiving images from the Simulator and to initiate the algorithm program to process the images.
[0085] Figure 5 In the illustrated embodiment, the test task splitting based on server topology enables the test to accurately adapt to the actual deployment environment, avoiding configuration chaos caused by unclear node division. Furthermore, the introduction of dependencies ensures that the execution order of process control instructions is consistent with the internal logic of the system, eliminating process crashes or functional failures caused by incorrect startup order. Simultaneously, combined with real-time process status monitoring, the system can dynamically detect anomalies, ensuring the stability and reliability of the testing process. It is understood that the testing process in this embodiment transforms from manual intervention to fully automated and orderly execution, shortening test preparation time and improving testing efficiency.
[0086] To improve the stability and reliability of the testing process and further reduce manual intervention, this application embodiment includes a corresponding restart mechanism for abnormal process states.
[0087] In some optional embodiments, if the process state of any process indicates that the process is running abnormally, the target process control instruction currently being executed is determined; according to the priority order, termination instructions are sent to the process corresponding to the target process control instruction and the process corresponding to the process control instructions executed before the target process control instruction; if the processes corresponding to all termination instructions successfully terminate, the step of executing process control instructions according to the priority order of multiple process control instructions and the process state of multiple processes is re-executed.
[0088] The target process control instruction can refer to a process control instruction that is currently being executed but has not yet been completed when an anomaly is detected. The process corresponding to this process control instruction is the process that experienced the anomaly.
[0089] For example, the process status can be obtained in each preset check cycle, and if the number of times the obtained running status that represents the abnormal running of the process is greater than a preset number threshold, the process is determined to be abnormal.
[0090] In one example, the preset check interval can be 5 seconds, and the preset check count threshold can be 3 times. That is, the process status is obtained every 5 seconds, and if the process status is detected as abnormal (such as error detection) 3 times consecutively, the process is determined to be abnormal.
[0091] For example, the process can automatically report its status, or the test center server can send a status acquisition instruction to obtain the process status. After receiving the status acquisition instruction, the process performs status detection and reports the process status.
[0092] In one example, the process management module implements real-time monitoring of the status of each process in the distributed simulation system using the following method, denoted as `check_process_by_name`, whose input parameter is the name of the process to be monitored. The process status monitoring mechanism can vary depending on the process's running location. Specifically: for processes running on the local server, the local process management module can directly read the operating system's ` / proc` filesystem to obtain process information, which contains detailed status data for all running processes in the system. For processes running on a remote server, the remote process management module uses a pre-defined communication channel (such as a communication channel built based on the SSH protocol) to execute the `ps` command to query the status of the process with the specified name and obtain the returned results. After obtaining the process status, the remote server can perform status judgment according to pre-defined logic. For example, if the process exists and its status is "running," the process is considered normal and a corresponding flag is returned; if the process does not exist, it is considered abnormally terminated and an exception flag is returned; if the process exists but its status is abnormal (such as zombie, uninterruptible, etc.), it is considered an abnormal process status and a corresponding flag is returned.
[0093] For example, a termination instruction can be used to control a process to stop running. It is understood that termination instructions can be sent in priority order to the process corresponding to the target process control instruction and to the process corresponding to the process control instruction executed before the target process control instruction.
[0094] In one example, processes can be shut down sequentially in the order of business services, application services, and basic services.
[0095] When a process is shut down, a graceful termination strategy can be employed. This strategy comprises at least two phases, each corresponding to a termination signal of varying strength. For example, upon initial confirmation of a process malfunction, a termination signal corresponding to the first phase is sent to the process. If the process remains open for a preset period, a termination signal corresponding to the second phase is sent. The termination signal in the first phase exerts less control over the process than the termination signal in the second phase.
[0096] In one example, the process management module terminates processes in a distributed simulation system using the following method, denoted as `terminate_by_name`. Its input parameters include the name of the process to be terminated and a configurable graceful timeout (default value 10 seconds). For instance, this method employs a two-phase termination strategy to ensure that processes have the opportunity for self-cleanup and can be forcibly terminated when necessary, avoiding resource leaks or process remnants. Specifically, the first phase is graceful termination. The process management module first sends a SIGTERM signal to the target process. This SIGTERM signal notifies the process that termination is required and allows the process to perform necessary cleanup operations, such as releasing resources, saving state, and closing network connections. After sending the SIGTERM signal, the process management module enters a waiting state, with a maximum waiting time of the preset graceful timeout (e.g., 10 seconds). During the waiting period, the process management module continuously monitors the process status, checking whether the process has exited normally. If the target process has not exited after the graceful timeout expires, it enters the second phase of forced termination. At this point, the process management module determines that graceful termination has failed and then enters the forced termination phase. The system sends a SIGKILL signal to the target process. This SIGKILL signal is handled directly by the operating system kernel and cannot be caught or ignored by the process, resulting in immediate forced termination. After termination, resource cleanup operations are performed to ensure that system resources such as file descriptors and memory associated with the process are properly released. Finally, the process management module returns the corresponding status based on the termination result, determining whether the process terminated successfully.
[0097] For example, if the process state of any process indicates an abnormal process execution, the currently executing target process control instruction is determined. Based on priority, termination instructions are sent to the process corresponding to the target process control instruction and the processes corresponding to process control instructions executed before the target process control instruction. If all processes corresponding to termination instructions successfully terminate, process control instructions continue to be executed according to the priority order of the multiple process control instructions and the process states of the multiple processes.
[0098] In one example, if the Simulator in the local controller is found to be in an abnormal state, the Simulator can be stopped first, and then DIFService can be stopped. Further, the processes can be restarted incrementally, starting DIFService first, and then the Simulator.
[0099] In another example, if the algorithm process that processes the simulation data in the remote controller is found to be in an abnormal state, it is determined that it is controlled by the MCN based on the dependency relationship. Therefore, it is possible to restart only the algorithm process that processes the simulation data and the MCN, without restarting the Simulator and DIFService in the local controller.
[0100] In another example, for faults that cannot be resolved by restarting the process, the fault type can be determined and the fault recovery strategy corresponding to the fault type can be used for fault recovery. Figure 6 This illustration shows a flowchart of a fault recovery process implemented through different fault recovery strategies, as provided in one embodiment of this application. Figure 6 As shown, in S610, the system initiates continuous process status monitoring to obtain the running status of each process in real time. In S620, when any abnormal process status is detected, fault diagnosis is performed. In S630, the fault type is determined based on the abnormal behavior, where the fault type includes at least four types: process abnormality, configuration abnormality, resource abnormality, and network abnormality.
[0101] In S641, if a process is determined to be abnormal, a process restart strategy is executed. This strategy involves first gracefully stopping the abnormal target process by sending a SIGTERM signal, waiting for the process to complete cleanup, clearing the system resources it occupied, and finally starting a new process instance. In S642, if a configuration error is determined, a configuration reload strategy is executed. This strategy includes: first verifying the validity of the current configuration; if verification fails, reverting to the default configuration; if verification succeeds, applying the new configuration and ensuring it takes effect. In S643, if a resource error is determined (e.g., memory leak or temporary file accumulation), a resource recovery strategy is executed. This strategy includes releasing memory occupied by the error and clearing useless temporary files. In S644, if a network error is determined, the system executes a network recovery strategy. This strategy includes: first testing the network connection; if the connection fails, attempting to rebuild the network connection; if the connection succeeds, confirming network recovery and re-establishing communication with relevant services.
[0102] like Figure 6 As shown, after any of the above strategies is executed, S650 is executed to verify the process status. If the verification passes, it indicates that the process has resumed normal operation, and S661 is executed to record the successful recovery log. If the verification fails, S662 is executed to implement the fault escalation handling mechanism. That is, the problem is reported to the alarm manager. The alarm manager will then handle the problem further according to preset rules, such as notifying maintenance personnel or attempting higher-level recovery methods.
[0103] In this embodiment, when an anomaly is detected, the system accurately identifies the affected scope based on priority order, terminating only processes that depend on the abnormal process. This avoids unnecessary global restarts, thereby preserving the normal operating process state to the greatest extent possible and reducing recovery time and resource waste. Furthermore, after all relevant processes have successfully terminated, the system re-executes the scheduling process according to a predetermined priority order, ensuring that restarted processes can start in an orderly manner following the correct dependencies, avoiding secondary failures caused by incorrect startup order.
[0104] In one optional implementation, to further improve the stability and real-time performance of the distributed simulation system test, a first communication channel and a second communication channel are established between the test center server and each server in the target simulation system. The first communication channel is used to transmit data related to the test task, and the second communication channel is used to transmit data related to server performance.
[0105] For example, the first communication channel can be built based on the Secure Shell (SSH) protocol, and the second communication channel can be built based on the Remote Direct Memory Access (RDMA) protocol.
[0106] In one example, Figure 7 This application illustrates a schematic diagram of the network protocol and communication layer architecture provided in one embodiment, as shown below. Figure 7As shown, the test control layer 710 in the local server includes a test controller 711, an SSH client 712, and an SFTP client 713, which are responsible for initiating test tasks, issuing commands, and transferring configuration files, respectively. The communication protocol layer 720 defines a variety of protocols to meet different communication needs, including: SSH protocol 721 (based on port 22) for remote command execution and process management; SSH file transfer protocol 722 (SFTP) for secure file transfer, such as sending YAML configuration files to various servers; gRPC remote procedure call protocol 723 (based on port 50001) for lightweight, high-performance inter-service communication, such as command interaction between scheduling services and algorithm services; and RDMA protocol 724 (based on InfiniBand network) for high-speed data transmission to meet the real-time requirements of large-scale data exchange during simulation. The network transport layer 730 includes a Transmission Control Protocol / Internet Protocol (TCP / IP) network 731, an InfiniBand network 732, and a separate management network 733, which respectively carry the data streams of the aforementioned protocols. The remote service layer 740 in the remote server deploys an algorithm service 741 (Algo_server), a scheduling service 742 (Schedule_server), and a monitoring service 743, which receive control commands through the underlying network and return execution results.
[0107] For example, the first communication channel is used to transmit data related to the test task, which may include target configuration file distribution, command sending, and test result feedback. The second communication channel is dedicated to transmitting server performance-related data, including monitoring indicators such as Central Processing Unit (CPU) load, RDMA bandwidth, and memory usage.
[0108] In this embodiment, the first communication channel is used to transmit data related to the test task, ensuring reliable transmission and timely response of test control commands; the second communication channel is dedicated to transmitting server performance-related data, ensuring that continuous performance data acquisition is not affected by test task execution. This separation design avoids data interaction during test task execution crowding out the performance monitoring channel, and also prevents performance data acquisition from interfering with the transmission of test task control commands, thereby ensuring the bidirectional stability and real-time performance of test control and performance monitoring. It is understood that using different communication channels enables accurate acquisition of hardware-level performance indicators even when the simulation system is running at full load, providing accurate and continuous data support for performance bottleneck analysis and fault diagnosis, and improving the reliability of distributed simulation system testing.
[0109] In some alternative embodiments, to evaluate the simulation system, server performance test data can be acquired via a second communication channel. Based on the performance test data and a preset performance index calculation algorithm, server performance indicators are determined, and based on these indicators and preset performance index thresholds, the server's performance test results are determined.
[0110] For example, performance monitoring data can be raw data collected from each server via a second communication channel. Performance monitoring data may include, but is not limited to, CPU load, memory usage, RDMA bandwidth, and disk input / output (I / O).
[0111] The preset performance index calculation algorithm can be a predefined rule used to convert raw performance test data into quantifiable and comparable performance indicators. Different performance indicators may have different corresponding performance index calculation algorithms.
[0112] In one example, a multi-dimensional performance monitoring module can be used to achieve real-time data acquisition and analysis of the operating status of a distributed simulation system. This performance monitoring module takes monitoring configuration parameters as input and outputs a real-time performance data stream, which can include both RDMA bandwidth monitoring and CPU load monitoring.
[0113] On one hand, RDMA bandwidth monitoring targets the InfiniBand network interface, with the data source being a hardware counter file in the Linux system, formatted as ` / sys / class / infiniband / {device} / ports / {port} / counters / `. Here, `{device}` represents the device name, and `{port}` represents the port number. The system creates an independent data acquisition thread and executes the `file_reader` function, which receives the device name, stop event, data list, and port number as parameters. Once the data acquisition thread enters a loop, it performs data acquisition once per second until the stop event is triggered. Data acquisition can be implemented as follows: construct the full path to the counter file (e.g., `port_xmit_data`), open and read the current counter value (in bytes), store the raw data in the data list, and set the acquisition frequency to 1 Hz. Further, the acquired raw byte data needs to be converted into bandwidth values. The module can obtain the byte difference by iterating through the data list and calculating the difference in the number of bytes between two adjacent samples. This byte difference is then multiplied by 8 to convert it to bits, multiplied by 4 times the transmission rate, and divided by 1e9 to finally obtain the real-time bandwidth value in gigabits per second (Gbps). This module supports a monitoring range of 0-100Gbps, with a data accuracy of 0.001Gbps and a sampling interval accurate to 1 second.
[0114] For example, Figure 8 A schematic diagram of the bandwidth calculation process provided in one embodiment of this application is shown, as follows: Figure 8 As shown, in S810, data is acquired through a data acquisition thread to obtain the raw byte sequence. Specifically, the hardware counter of the InfiniBand network interface can be sampled at fixed time intervals (e.g., 1 second) to obtain a series of raw byte data sequences arranged in time order [t0, t1, t2, t3, ...], denoted as [C0, C1, C2, C3, ...]. Each data point in the raw byte data sequence represents the cumulative number of bytes transmitted up to that sampling time.
[0115] In S820, the incremental difference of the number of bytes between adjacent sampling points in the original byte sequence is calculated to obtain the byte increment sequence. The incremental difference can be calculated using the following formula (1): ΔCi represents the actual number of bytes transmitted within a 1-second time interval, and C(i+1) and C(i) represent the number of bytes acquired in the (i+1)th second and the i-th second, respectively. For example, the original byte data sequence [1000, 1500, 2000, 2800, 3600, 4400, ...] can be represented as [500, 500, 800, 800, 800, ...] after increment difference calculation.
[0116] In S830, the bandwidth sequence is obtained by bandwidth conversion of the byte increment sequence according to the bandwidth calculation formula. The bandwidth calculation formula can be constructed based on the RDMA transmission characteristics and can be expressed as the following formula (2): bandwidth is the bandwidth value. It can be understood that the bandwidth calculation formula shown in formula (2) is used to convert the byte difference into a bandwidth value in Gbps. Among them, multiplying by 8 converts the byte to bits, multiplying by 4 indicates that the transmission rate is considered by 4 times (corresponding to the transmission characteristics of the InfiniBand network), and dividing by 1e9 converts the result to Gbps. Continuing with the previous example, the bandwidth sequence after the byte increment sequence [500, 500, 800, 800, 800, ...] is converted can be expressed as [0.016, 0.016, 0.026, 0.026, 0.026 Gbps].
[0117] In S840, after obtaining the bandwidth sequence, statistical analysis is performed on the bandwidth sequence to extract key statistical features. These key statistical features may include the maximum value (0.026 Gbps), minimum value (0.016 Gbps), and average value (0.022 Gbps). Simultaneously, a filtering threshold (e.g., less than 0.001 Gbps) can be set to remove noisy data or identify idle periods.
[0118] On the other hand, CPU load monitoring targets the CPU load of both local and remote servers. For local servers, the `uptime` command is executed to obtain system load information. The "loadaverage" field is parsed from the `uptime` command output to extract the average load values over 1 minute, 5 minutes, and 15 minutes. The parsing process includes: executing the command, obtaining standard output, extracting the load string, splitting the string, and converting it to a floating-point number. For remote servers, the system connects to the target server via the SSH protocol and executes the `exec_remote_cmd` command. The `exec_remote_cmd` method accepts the hostname, username, and the `uptime` command as parameters. After establishing the SSH connection, the command is executed and the output is obtained. After disconnecting, the data is processed using the same parsing logic as local monitoring. The network latency for remote monitoring is less than 100 milliseconds. The CPU load is collected every 10 seconds, and the monitoring range covers load rates from 0 to the number of server cores.
[0119] In yet another example, Figure 9 This application illustrates a schematic diagram of a data processing flow according to an embodiment of the present application, such as... Figure 9 As shown, the performance monitoring module's architecture, from bottom to top, consists of a data source layer (910), a data acquisition layer (920), a data processing layer (930), an analysis and calculation layer (940), and a result output layer (950). These layers work together to complete the entire process from acquiring raw data to presenting the final results. Specifically, the data source layer (910) is located at the bottom of the architecture and forms the data foundation for performance monitoring. This layer contains various types of raw data sources, including: an RDMA hardware counter (911), an uptime command (912), a process log file (913), and a network statistics interface (914). The RDMA hardware counter provides byte-level data transfer data for the InfiniBand network; the system uptime command provides CPU load information; the process log file records the running status and event timestamps of each process under test; and the network statistics interface provides metrics such as network throughput and connection count.
[0120] The data acquisition layer 920 is responsible for obtaining raw data from various data sources. This layer can achieve efficient data acquisition through a multi-threaded concurrent acquisition mechanism. For example... Figure 9 As shown, RDMA acquisition thread 921 reads the hardware counter file at a fixed frequency; CPU acquisition thread 922 obtains system load by executing the uptime command or via remote SSH; log parsing thread 923 monitors and analyzes process log files in real time. Network acquisition thread 924 calls the system network interface to obtain statistical information. Each acquisition thread runs independently and does not interfere with the others.
[0121] Data processing layer 930 is used for preliminary cleaning and preprocessing of the collected raw data. This layer performs three operations: outlier filtering, removing invalid data caused by sampling jitter or transient system anomalies; data aggregation, merging high-frequency sampled data according to time windows to reduce subsequent computational burden; and time series alignment, unifying data from different data sources and sampling frequencies onto the same time coordinate system, laying the foundation for multi-indicator correlation analysis.
[0122] The analysis and computation layer 940 is used for in-depth calculations and metric extraction of the processed data. This layer can be used for multi-dimensional analysis and calculation of the data, such as: bandwidth calculation: converting the difference of the RDMA hardware counter into a real-time bandwidth value; load statistics: performing statistical calculations such as average and peak CPU load data; transmission time analysis: calculating the transmission latency of data between different nodes by combining timestamps in the process log; and utilization calculation: comprehensively considering multi-dimensional data to assess the actual utilization rate of system resources.
[0123] The results output layer 950 can present analysis and calculation results to technical personnel in multiple formats. This layer supports diverse output methods, such as: providing structured performance data tables in comma-separated values (CSV) reports for easy offline analysis and archiving; displaying metric trends visually in performance charts to help quickly identify anomalies; persistently storing critical events and alarm information using log archiving for easy problem backtracking; and dynamically refreshing the latest performance data through a real-time monitoring panel for real-time observation during testing.
[0124] Furthermore, server performance metrics can be analyzed to generate performance analysis reports.
[0125] For example, multi-metric correlation analysis can be performed. For instance, when RDMA bandwidth suddenly increases, the CPU load can be simultaneously assessed to identify potential performance bottlenecks. Bandwidth anomaly detection can be triggered by preset anomaly detection thresholds, such as triggering a load warning when the CPU load exceeds 80% for three consecutive minutes, or triggering bandwidth anomaly detection when the RDMA bandwidth deviates from the historical baseline by more than 30%. Furthermore, a tiered strategy can be adopted for data storage: real-time data is stored in a memory buffer, retaining a maximum of the most recent 1000 data points for fast access. Historical data is periodically persisted to a CSV file for offline analysis. Alarm data is recorded separately in an independent log file for easy problem tracking.
[0126] Furthermore, the performance monitoring module can reduce performance overhead through relevant parameter settings, such as: setting the RDMA data acquisition frequency to 1Hz and the CPU data acquisition frequency to 0.1Hz; ensuring memory usage is less than 50MB under full-indicator monitoring; reducing the CPU overhead of the acquisition process itself to less than 1%; and reducing the network overhead of monitoring data transmission to less than 1Mbps. This ensures that performance monitoring itself does not cause significant interference to the system under test, achieving efficient and accurate real-time observation capabilities for the distributed simulation system.
[0127] In addition, the performance monitoring module can analyze the server's performance metrics to identify performance bottlenecks; and based on the identified performance bottlenecks, determine server performance optimization strategies.
[0128] In one example, Figure 10 This illustration shows a flowchart of a performance optimization strategy determination provided in one embodiment of this application, as follows: Figure 10 As shown, in S1010, performance metrics are collected from each server, including CPU utilization, memory utilization, network I / O throughput, RDMA bandwidth, and service response time.
[0129] In S1020, performance metrics are analyzed to obtain bottleneck analysis results. For example, current performance metrics can be compared with historical baselines or preset thresholds to identify bottleneck types. Specifically, this includes: CPU bottleneck detection, such as determining if computing resources are saturated; memory bottleneck detection, such as identifying memory leaks or frequent swapping; network bottleneck detection, such as analyzing congestion on traditional network interfaces; and RDMA bottleneck detection, such as evaluating the transmission efficiency of high-speed interconnect links. Through correlation analysis of multiple metrics, the system can accurately pinpoint the key factors limiting system performance.
[0130] In the S1030, corresponding optimization schemes are matched based on bottleneck analysis results. The system has a built-in strategy library for different bottleneck types. For example, for CPU bottlenecks, thread pool tuning strategies can be selected to adjust the number of concurrent threads to balance the load. For memory bottlenecks, cache tuning strategies are used to optimize the data buffering mechanism. For network bottlenecks, connection pool tuning strategies are applied to reduce connection establishment overhead. For RDMA bottlenecks, queue depth tuning strategies are implemented to adjust data transmission queue parameters to match hardware capabilities. The optimization strategy selection process comprehensively considers the current system load, hardware configuration, and test objectives to ensure the effectiveness of the selected strategies.
[0131] In S1040, the determined optimization strategy is translated into parameter adjustment instructions and applied to the corresponding objects. This includes a feedback mechanism to continuously monitor system performance and verify the effects after optimization adjustments. By comparing changes in key metrics before and after optimization, the system automatically evaluates the optimization effect. If the expected results are not achieved, the bottleneck analysis process is retried, forming a new round of optimization iterations; if the expected results are achieved, the optimized parameter combination is recorded in the strategy library for reference in subsequent testing. It is understandable that a cyclical monitoring approach can be used to optimize server performance.
[0132] In this embodiment, a separately configured second communication channel ensures the stability and real-time performance of data acquisition, avoiding contention for network resources with test task control commands and ensuring that monitoring data can be continuously and accurately transmitted to the test center even when the simulation system is running at full load. Furthermore, a preset performance index calculation algorithm converts raw hardware counter readings and system status data into performance indicators with clear physical meaning, enabling testers to intuitively understand the system's operating status. Moreover, by comparing with preset thresholds, the system can automatically identify performance anomalies and output standardized test results.
[0133] Below, in conjunction with Figures 11 to 17 The following examples illustrate the simulation testing method.
[0134] in, Figure 11 A schematic diagram of the architecture of an automated testing platform for a distributed simulation system provided in one embodiment of this application is shown; Figure 12 This paper illustrates a communication network topology diagram of a distributed simulation system provided in one embodiment of the present application. Figure 13 This illustration shows a flowchart of the dynamic generation of a target configuration file according to an embodiment of this application; Figure 14 This illustration shows a schematic diagram of the content composition of a JSON configuration file provided in one embodiment of this application; Figure 15 This illustration shows a schematic diagram of the content composition of a YAML configuration file provided in one embodiment of this application; Figure 16 This illustration shows a schematic diagram of the content composition of a dynamic configuration script provided in one embodiment of this application; Figure 17 A timing diagram of concurrent execution of a target configuration file provided in one embodiment of this application is shown.
[0135] like Figure 11As shown, the test platform 1100 of the distributed simulation system adopts a layered architecture design, including three layers: the test control center 1110, the network communication layer 1120, and the remote server layer 1130. The test control center 1110 (Master) is deployed on the local server and is responsible for the unified scheduling and management of the entire test process. The test control center contains four core functional modules: a configuration generation module 1111, a process management module 1112, a performance monitoring module 1113, and a result analysis module 1114. The configuration generation module 1111 is responsible for reading the initial configuration file in JSON format and converting it into multiple target configuration files in YAML format. The process management module 1112 is responsible for the lifecycle management of various processes on the local and remote servers, including starting, monitoring, and terminating them. The performance monitoring module 1113 is responsible for collecting server performance data, including indicators such as RDMA bandwidth and CPU load. The result analysis module 1114 is responsible for processing the data collected during the test and generating test reports.
[0136] And, as Figure 11 As shown, the network communication layer 1123 includes two types of communication channels: RDMA high-speed network 1121 is used to carry large-scale data transmission during the operation of the simulation system, meeting the requirements of low latency and high throughput; SSH / SFTP communication channel 1122 is used for the issuance of control commands, transmission of configuration files and return of status information between the test control center and the remote server.
[0137] Remote server layer 1130 contains multiple remote servers 1131. Figure 11 The diagram shows two remote servers (IP addresses 172.16.19.1 and 172.16.19.2 respectively) and a local server 1132 (IP address 127.0.0.1). Each remote server runs an Algo Server Process, an RDMA Monitor module, and a system monitoring module. The local server runs a Schedule service, a Simulator, and a local monitoring module.
[0138] like Figure 12 As shown, the test platform uses an independent management network to interconnect all components. The communication network topology includes management network 1210. Management network 1210 corresponds to network segment 172.16.19.0 / 24 and is a dedicated transmission link for system control data. It connects control server 1221, first algorithm (algo_server) server 1222, and second algorithm server 1223, respectively, carrying the interaction of test control commands, configuration files, node status, and monitoring data. It is the underlying physical network carrier for the SSH / SFTP communication channel.
[0139] The communication network topology also includes a distributed server node layer 1220, comprising the aforementioned control server 1221, first algorithm server 1222, and second algorithm server 1223. Control server 1221 is a local server with IP address 127.0.0.1, internally deploying the pytest testing framework, Schedule_server scheduling service, Simulator simulation simulator, local monitoring module, and RDMA module, responsible for the entire test process scheduling, simulation scenario construction, and local status monitoring. First algorithm server 1222 and second algorithm server 1223 are two different remote servers with identical deployment architectures, both containing the algo_server algorithm service process, RDMA module, and system monitoring module, responsible for algorithm processing of simulation data, high-speed data interaction, and node performance data acquisition.
[0140] Among them, an independent RDMA high-speed network 1230 is set up between the distributed server node layers. It is built based on InfiniBand technology and connects to the RDMA module of each server to carry the data transmission of simulation services. It is physically isolated from the management network to avoid the service data flow and control command flow competing for bandwidth.
[0141] In addition, the communication network topology also includes a locally shared storage system 1240. This system has two core storage paths, / mnt / share / data / image and / mnt / mass / DFIService, which communicate with the control server 1211, the Schedule_server scheduling service, the Simulator simulation simulator, and the algo_server algorithm service processes on two remote servers to provide a unified shared storage service for test source data, configuration files, and result data.
[0142] Furthermore, such as Figure 13 As shown, in S1310, the JSON configuration file is read and parsed to obtain the test task list. The test task list is defined in JSON data format and includes simulation devices and simulators. Specifically, as... Figure 14As shown, the JSON configuration file 1400 includes parameters for the simulation device defined by the jobs node 1410 and parameters for the simulator defined by the simulator node 1420. For example, the jobs node may include two sets of parameters: 8bit_20x_SPA and 12bit_40x_bare, which define parameters for different simulation devices. The simulator node may include a hosts [0] sub-node 1421, which specifies the network parameters of the simulator, including the service port port: 50001 and the host IP address ip: 127.0.0.1. It can be understood that the simulation device can represent the front-end device to be simulated, such as a camera, etc.
[0143] In S1320, the data in the test task list is reconstructed according to the YAML data structure to obtain the YAML configuration file. It is understandable that if the simulation device needs to be simulated, the simulator must first provide a stable operating environment and data output carrier. Therefore, the simulator node in the YAML configuration file corresponds to a higher level. Specifically, the hosts [0] node 1503 is set under the simulator node 1502 in the YAML configuration file 1501. It is directly associated with the hosts [0] in the JSON configuration file and inherits the network parameters of the simulator defined in the JSON configuration file, namely port:50001 and ip: 127.0.0.1, to ensure that the simulator can start and communicate normally on the specified server node. At the same time, the parameters of the simulation device defined in the JSON configuration file (i.e., 8bit_20x_SPA, 12bit_40x_bare) are filled into the cameras sub-node 1504 set in the YAML configuration file to simulate the camera and output image data.
[0144] In S1330, data from the YAML configuration file is passed to the dynamic configuration script via parameterized mapping rules to complete the parameterized entry point configuration for test execution. The dynamic configuration script splits the hierarchical structure of the YAML configuration file into four executable configuration branches. Specifically, as follows... Figure 16As shown, the dynamic configuration script 1600 includes a scheduling configuration 1610 (Schedule), which defines the list of remote servers participating in the test and clarifies the range of schedulable hardware computing resources; an algorithm service configuration 1620 (Algo_Server), which sets running parameters for each remote server node, including allocating the corresponding number of CPU cores and worker threads to slave nodes 1621 such as slave1 and slave2, achieving fine-grained allocation of distributed computing resources. It also includes a server count configuration 1630, used to preset a set of adaptable test cluster size values to cover test scenarios with different numbers of nodes; and a thread count configuration 1640, used to preset a set of adaptable concurrent processing pressure values to cover performance test scenarios under different loads. It can be understood that the dynamic configuration script generated through the above parameterized mapping can directly achieve full lifecycle control of all processes related to the entire test process, including the startup, status monitoring, and abnormal termination of the scheduling service process and simulation simulator process on the local server, as well as the remote scheduling, resource configuration, and operation management of the algorithm service process on the remote server.
[0145] Furthermore, YAML files can be processed concurrently. For example... Figure 17 As shown, Figure 17 The five participating entities are arranged horizontally: main thread 1701, first test thread 1702 (responsible for the 8×1 single-server test scenario), second test thread 1703 (responsible for the 16×1 single-server test scenario), first remote server 1704, and second remote server 1705. The vertical axis is the timeline, showing the message interaction sequence from top to bottom.
[0146] During the test initiation phase, the main thread concurrently launches the first and second test threads. The first test thread begins processing the 8×1 single-server test scenario: it first generates the corresponding test configuration for this scenario, uploads the configuration file to remote server 1, then starts the process under test on the first remote server and begins performance monitoring, collecting performance data after the test execution is complete. Simultaneously, the second test thread processes the 16×1 single-server test scenario in parallel: it similarly executes the process of generating configuration, uploading configuration, starting the process, starting monitoring, executing the test, and collecting data, uploading its configuration to the second remote server. This process overlaps with the first test thread's operations in time. After both single-server test scenarios are completed, they each return a completion status to the main thread. The dual-server test phase then begins, with the main thread launching the 8×2 test scenario. The main thread simultaneously uploads the test configuration to both the first and second remote servers. Both servers launch their respective MCN processes, then execute the test in parallel and collect data. The test execution processes on both servers overlap in time, without waiting for each other, and each returns the test results to the main thread upon completion. Understandably, through the aforementioned multi-threaded concurrency and parallel execution mechanisms, the testing platform can run multiple test scenarios simultaneously, shortening the overall testing time and improving testing efficiency.
[0147] Furthermore, the overall testing process is as follows: Figure 18 As shown, in S1801, the test environment is initialized. The system starts the test control center, establishes communication connections with each remote server, and checks the reachability and basic environment configuration of each server.
[0148] In S1802, process status is checked. The system checks the current status of the processes under test on each server through the process management module. If an abnormal process status is detected, the process restart procedure is initiated: the scheduling service (ScheduleService), the simulation simulator (Simulator), and the remote algorithm service (Algo_Service) are stopped sequentially. Then, the scheduling service is restarted, and a 15-second wait is waited to ensure it is fully started before starting the remote algorithm service. Another 18-second wait is then waited before starting the simulation simulator. It is understandable that since the test environment may not start from a zero state, checking the process status serves as a way to clean up the test environment.
[0149] In S1803, the test configuration is generated. The configuration generation module generates the configuration file required for this test based on the preset test parameters.
[0150] In S1804, start the simulation simulator. Start the simulation simulator process on the local server to prepare for the simulation task.
[0151] In S1805, start performance monitoring. The performance monitoring module begins collecting performance data from each server, including RDMA bandwidth, CPU load, etc.
[0152] In S1806, test cases are executed. The system executes test cases according to preset test scenarios, driving the simulation system under test to run.
[0153] In S1807, the system waits for the test to complete and performs a timeout check. The system continuously waits for the test to complete while simultaneously executing S1808 to monitor the test execution time. If the test completes within the specified time, S1809 is executed; otherwise, S1810 is executed to abort the current test.
[0154] In S1809, stop monitoring. After the test is complete, stop data collection from the performance monitoring module.
[0155] In S1811, monitoring data is collected. The system collects performance data gathered throughout the testing process, preparing it for analysis.
[0156] In S1812, performance metrics are calculated. The collected raw data is processed to calculate key performance indicators such as bandwidth, load, and response time.
[0157] In S1813, a test report is generated. The results analysis module generates a structured test report based on the calculated performance metrics.
[0158] In S1814, save the test results. Persist in storing the test report and related log files for later analysis and traceability.
[0159] Furthermore, this embodiment encapsulates the core functionality of the testing platform into multiple cooperating classes, forming a clear software architecture. Specifically, the testing platform may include the RemoteProcessManager class, ConfigGenerator class, PerformanceMonitor class, TestCase class, and TestResult class.
[0160] The `RemoteProcessManager` class is responsible for process management on a remote server. This class includes attributes such as hostname, username, password, key file path, and port number, as well as an SSH client object and a connection status indicator (connected). It uses `connect()` to establish an SSH connection, `disconnect()` to disconnect, `start_program()` to start a program on the remote server and return the process ID, `check_process_by_name()` to check the process status based on the process name, and `terminate_by_name()` to terminate the process based on the process name with a specified graceful timeout. This class supports SSH key authentication, providing secure and reliable remote process management capabilities.
[0161] The `ConfigGenerator` class is responsible for automatically generating test configurations. This class implements JSON to YAML format conversion and supports parameterized configuration. Specifically, it can generate a list of configurations for multiple jobs from a JSON file using `create_job_conf()`, generate a configuration file for a single job using `create_singleJob_conf()`, serialize the configuration dictionary into a YAML file using `generate_yaml_config()`, and parse the JSON configuration file and return a configuration dictionary using `parse_json_config()`.
[0162] The `PerformanceMonitor` class is responsible for collecting and calculating performance data. This class contains a stop event flag (`stop_flag`), a data list dictionary (`datalist`), a collection thread dictionary (`threads`), and an RDMA bandwidth data dictionary (`rdma_bandwidth`). `start_monitoring()` can be used to start performance monitoring of a specified server, `stop_monitoring()` to stop monitoring, `collect_rdma_data()` is the execution function for the RDMA data collection thread, `calculate_bandwidth()` calculates the bandwidth value based on the raw byte data, and `monitor_cpu_load()` collects CPU load data.
[0163] The `TestCase` class is the main class for test cases, responsible for managing the entire test lifecycle. This class contains attributes such as configuration name (`config_name`), server list (`servers`), thread count (`thread_num`), and server count (`server_num`). It can use `setup()` for environment preparation before testing, `run_test()` for executing tests, `teardown()` for post-test cleanup, and `collect_results()` for collecting test results.
[0164] The `TestResult` class is responsible for processing test results and generating reports. This class contains attributes such as test case name (`case_name`), run time (`run_time`), number of jobs (`job_count`), number of defects (`defect_count`), a performance metric dictionary (`performance_metrics`), an RDMA metric dictionary (`rdma_metrics`), and a CPU metric dictionary (`cpu_metrics`). Specifically, `calculate_metrics()` is used to calculate various performance metrics, `generate_csv_report()` is used to generate a CSV format test report, and `save_log_files()` is used to save log files.
[0165] It is understandable that, through the collaboration of the above-mentioned classes, the testing platform in this application embodiment achieves full-process automation from configuration generation, process management, performance monitoring to result analysis. Each module has clear responsibilities and well-defined interfaces, and has good scalability and maintainability.
[0166] Based on the simulation testing method, this application also provides specific embodiments of the testing device for the simulation system.
[0167] like Figure 19 As shown, the simulation testing device 1900 provided in this application embodiment includes: an acquisition module 1901, a determination module 1902, a construction module 1903, and a scheduling module 1904.
[0168] Module 1901 is used to obtain task information for test tasks; The determination module 1902 is used to determine the topology of the target simulation system that matches the task information and the task type based on the preset topology of different simulation systems corresponding to different task types. Module 1903 is used to build dependencies between processes that execute test tasks based on the topology of the target simulation system. The processes are processes in the server of the target simulation system. The scheduling module 1904 is used to schedule processes to execute test tasks based on dependencies and process states.
[0169] As an optional embodiment, the scheduling module 1904 schedules processes to execute test tasks based on dependencies and process states in the following manner: a splitting unit, used to split the test task according to the topology of the target simulation system to obtain at least one test subtask corresponding to the server of the target simulation system; a generation unit, used to generate process control instructions corresponding to multiple processes for executing the test subtasks according to dependencies and test subtasks; a determination unit, used to determine the priority order of multiple process control instructions according to dependencies; and an execution unit, used to execute the process control instructions according to the priority order of multiple process control instructions and the process states of multiple processes.
[0170] As an optional embodiment, the process includes: a main process deployed in the test center server for providing basic services and a simulation process for generating simulation data; a control process deployed in at least one remote server for communication and scheduling control and an algorithm process for processing the simulation data; multiple process control instructions include: a first process control instruction for starting the main process, a second process control instruction for starting the simulation process, a third process control instruction for starting the control process, and a fourth process control instruction for starting the algorithm process; the determining unit determines the priority order of the multiple process control instructions according to the dependency relationship in the following manner: according to the dependency relationship, the priority order of the multiple process control instructions is determined from high to low as: the first process control instruction, the second process control instruction, the third process control instruction, and the fourth process control instruction.
[0171] As an optional embodiment, during the execution of process control instructions according to the priority order of multiple process control instructions and the process states of multiple processes, the apparatus further includes: a second determining unit, configured to determine the currently executing target process control instruction when the process state of any process indicates an abnormal process operation; a sending unit, configured to send termination instructions to the process corresponding to the target process control instruction and the process corresponding to the process control instructions executed before the target process control instruction according to the priority order; and an execution unit further configured to re-execute the step of executing process control instructions according to the priority order of multiple process control instructions and the process states of multiple processes when all processes corresponding to termination instructions have successfully terminated.
[0172] As an optional embodiment, a first communication channel and a second communication channel are set up between the test center server and each server in the target simulation system. The first communication channel is used to transmit data related to the test task, and the second communication channel is used to transmit data related to server performance.
[0173] As an optional embodiment, the apparatus further includes: an acquisition module, configured to acquire server performance test data through a second communication channel; a performance index determination module, configured to determine server performance indexes based on the performance test data and a preset performance index calculation algorithm; and a test result determination module, configured to determine server performance test results based on server performance indexes and a preset performance index threshold.
[0174] As an optional embodiment, the acquisition module 1901 acquires the task information of the test task in the following manner: a receiving unit, used to receive the initial configuration file of the test task and parse the initial configuration file to obtain a test task list containing at least one test task, wherein the test task list contains the initial test task information in a preset data format corresponding to the test task; a conversion unit, used to convert the initial test task information corresponding to the test task in the test task list according to the target data structure to obtain the target configuration file corresponding to the test task; and an acquisition unit, used to acquire the task information of the test task by reading the target configuration file.
[0175] Based on the simulation testing method, this application also provides specific embodiments of the simulation testing equipment.
[0176] Figure 20 A schematic diagram of the hardware structure of the simulation test equipment provided in an embodiment of this application is shown.
[0177] The simulation test device 2000 may include a processor 2001 and a memory 2002 storing computer program instructions.
[0178] Specifically, the processor 2001 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0179] Memory 2002 may include mass storage for data or instructions. For example, and not limitingly, memory 2002 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 2002 may include removable or non-removable (or fixed) media. Where appropriate, memory 2002 may be internal or external to a device. In a particular embodiment, memory 2002 is a non-volatile solid-state memory.
[0180] The processor 2001 reads and executes computer program instructions stored in the memory 2002 to implement any of the simulation test methods in the above embodiments.
[0181] In one example, the electronic device may also include a communication interface 2003 and a bus 2010. Wherein, as... Figure 20 As shown, the processor 2001, memory 2002, and communication interface 2003 are connected through bus 2010 and complete communication with each other.
[0182] The communication interface 2003 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0183] Bus 2010 includes hardware, software, or both, that couples the components of the electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 2010 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0184] Furthermore, in conjunction with the image defect classification method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the simulation testing methods in the above embodiments.
[0185] In addition, in conjunction with the simulation testing methods in the above embodiments, this application embodiment can provide a computer program product for implementation. When the instructions in the computer program product are executed by the processor of an electronic device, the electronic device performs the simulation testing method provided by any aspect of the above embodiments of this application.
[0186] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0187] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0188] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0189] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0190] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method of emulation testing, characterized by, Applied to a test center server, the method includes: Obtain task information for the test task; Based on the preset topology of different simulation systems corresponding to different task types, determine the topology of the target simulation system that the task information corresponds to the task type. Based on the topology of the target simulation system, a dependency relationship between processes for executing the test task is constructed, wherein the processes are processes in the server of the target simulation system; Based on the dependencies and the process state of the process, the process is scheduled to execute the test task.
2. The emulation test method of claim 1, wherein, The step of scheduling the process to execute the test task based on the dependency relationship and the process state of the process includes: Based on the topology of the target simulation system, the test task is split to obtain at least one test subtask corresponding to the server of the target simulation system. Based on the dependencies and the test subtasks, process control instructions are generated for each of the multiple processes, which are processes used to execute the test subtasks. Based on the dependencies, the priority order of the multiple process control instructions is determined; The process control instructions are executed according to their priority order and the process states of the plurality of processes.
3. The emulation test method of claim 2, wherein, The process includes: a main process deployed in the test center server to provide basic services and a simulation process to generate simulation data; and a control process deployed in at least one remote server for communication and scheduling control and an algorithm process for processing simulation data; The plurality of process control instructions include: a first process control instruction for starting the overall process, a second process control instruction for starting the simulation process, a third process control instruction for starting the control process, and a fourth process control instruction for starting the algorithm process; Based on the dependencies, the priority order of the multiple process control instructions is determined, including: Based on the dependency relationship, the priority order of the multiple process control instructions is determined from high to low as follows: the first process control instruction, the second process control instruction, the third process control instruction, and the fourth process control instruction.
4. The emulation test method of claim 2, wherein, In the process of executing the process control instructions according to the priority order of the plurality of process control instructions and the process states of the plurality of processes, the method further includes: If the process state of any of the aforementioned processes indicates an abnormal process operation, determine the target process control instruction currently being executed; According to the priority order, a termination command is sent to the process corresponding to the target process control instruction and the process corresponding to the process control instruction executed before the target process control instruction. If all processes corresponding to the termination instructions successfully terminate, the step of executing the process control instructions according to the priority order of the multiple process control instructions and the process states of the multiple processes is re-executed.
5. The emulation test method of claim 1, wherein, The test center server and each server in the target simulation system are provided with a first communication channel and a second communication channel. The first communication channel is used to transmit data related to the test task, and the second communication channel is used to transmit data related to server performance.
6. The emulation test method of claim 5, wherein, The method includes: The server's performance testing data is obtained through the second communication channel; The performance indicators of the server are determined based on the performance test data and the preset performance indicator calculation algorithm. The performance test results of the server are determined based on the server's performance metrics and preset performance metric thresholds.
7. The emulation test method of claim 1, wherein, The process of obtaining the task information for the test task includes: Receive the initial configuration file of the test task and parse the initial configuration file to obtain a test task list containing at least one test task, wherein the test task contains initial test task information in a preset data format; The initial test task information corresponding to the test tasks in the test task list is transformed according to the target data structure to obtain the target configuration file corresponding to the test task; The task information of the test task is obtained by reading the target configuration file.
8. An electronic device, comprising: The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the simulation test method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the simulation testing method as described in any one of claims 1-7.
10. A computer program product, characterised in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the simulation test method as described in any one of claims 1-7.