An automated testing method and system for malicious applications of an ARM server

CN122845286APending Publication Date: 2026-09-29四川华鲲振宇智能科技有限责任公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611293686.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

现有方案对虚拟化资源的管理能力难以适配ARM虚拟机集群场景,不支持测试任务的批量并行调度,规模化测试的执行效率有限

Benefits of technology

(1)通过构建适配ARM架构的虚拟化测试环境,串联样本管理、任务调度、执行控制与流量分析完整流程,实现ARM服务器场景下恶意应用动态测试的自动化有序开展;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845286A_ABST
    Figure CN122845286A_ABST
Patent Text Reader

Abstract

The application discloses an automatic testing method and system for malicious applications of ARM servers, and belongs to the technical field of information security. The method receives a malicious application sample file and platform type information, calculates a sample unique identifier, and stores corresponding metadata information; receives test task definition information and target virtual machine selection information containing the sample unique identifier, completes virtual machine running state verification and network resource association; generates an automatic arrangement script and distributes it to the target virtual machine, matches corresponding operation execution logic to complete a test operation sequence; captures network flow data in the target virtual machine, analyzes and generates structured analysis data, and completes storage. The scheme can adapt to ARM architecture server scenarios to carry out dynamic testing of malicious applications, improve the reuse degree of cross-platform test scripts, enrich the dimension of network behavior analysis, and realize automatic and large-scale testing of malicious applications of ARM servers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information security technology, and in particular to an automated testing method and system for malicious applications targeting ARM servers. Background Technology

[0002] With the development of cloud computing and edge computing technologies, ARM architecture servers, with their low power consumption and high concurrency, are increasingly being used in data centers and various specific computing scenarios. The corresponding server operating system ecosystem is constantly improving, and the types of business and data volume they support are gradually increasing. At the same time, malicious application threats targeting server operating environments are constantly evolving. Various malicious codes, such as cryptocurrency mining programs, ransomware, and backdoors, are impacting the secure and stable operation of servers. The analysis and testing of malicious applications has become an important research direction in the field of information security. Currently, automated testing technology for malicious applications is under continuous development. Dynamic analysis solutions based on virtualization sandboxes are widely used in the industry. These solutions typically combine virtualization technology to build isolated test environments, use automated control tools to simulate application operating conditions, and integrate network traffic collection and behavior monitoring capabilities to complete the capture and analysis of malicious application behavior. Mainstream malicious application testing systems are mostly built around traditional general-purpose computing architectures, forming relatively mature sample management, task scheduling, and behavior analysis processes that can support the batch testing needs of malicious applications in common scenarios.

[0003] Existing automated testing solutions for malicious applications have several limitations when adapted to ARM server scenarios, making it difficult to meet the testing requirements of this environment. The virtualization underlying of existing testing systems is mostly designed for traditional computing architectures, making it impossible to directly run malicious application samples on the ARM architecture. This hinders the construction of a native ARM server testing environment, preventing effective dynamic analysis of such samples. Existing automated control solutions are not uniformly adapted to the ARM server operating environment. Different operating system distributions require the development of separate control scripts, and there is a lack of unified calling interfaces between different technical approaches. This results in low reusability of test scripts and high system usage and maintenance costs. Existing solutions have poor virtualization resource management capabilities for ARM virtual machine cluster scenarios, do not support batch parallel scheduling of test tasks, and have limited execution efficiency for large-scale testing. Furthermore, existing traffic acquisition solutions often use external bypass deployment, failing to fully capture network communication data within the virtual machine. Their ability to capture hidden communication scenarios within the server is insufficient, making it difficult to support comprehensive network behavior analysis of malicious applications on ARM servers. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an automated testing method and system for malicious applications targeting ARM servers.

[0005] The objective of this invention is achieved through the following technical solution: An automated testing method for malicious applications targeting ARM servers is provided, which includes the following steps: S1. Receive malicious application sample files and platform type information, calculate the hash value of the malicious application sample files to generate a unique sample identifier, store the malicious application sample files, and record the metadata information corresponding to the malicious application sample files. S2. Receive test task definition information and target virtual machine selection information containing the unique identifier of the sample, generate a unique identifier of the test task, confirm the running status of the virtual machine, perform a startup operation on virtual machines that are not running, detect the network connection status of the virtual machine, and associate the network address information of the virtual machine that passes the detection with the unique identifier of the test task. S3. Generate an automated orchestration script based on the test task definition information, distribute the automated orchestration script and malicious application sample file to the target virtual machine, match the corresponding operation execution logic, and execute the operation sequence corresponding to the automated orchestration script; S4. Capture network traffic data inside the target virtual machine during the execution of the malicious application sample file, collect test execution result data and network traffic data, parse the network traffic data to generate structured analysis data, and store the structured analysis data and test execution result data.

[0006] Furthermore, step S1 includes the following sub-steps: S1.1. Receive malicious application sample files and platform type selection information, parse the file data and form data in the upload request, and extract the binary content of the malicious application sample files and platform type selection information; S1.2. Perform a hash operation on the binary content of the malicious application sample file to generate a unique sample identifier for unique identification of the sample; S1.3. Store the malicious application sample files to the corresponding storage locations according to the set path rules, and record the storage path information corresponding to the malicious application sample files; S1.4. Summarize the metadata information of malicious application sample files. The metadata information includes the sample's unique identifier, file size information, upload time information, platform type information, and storage path information. Write the metadata information into the database storage medium.

[0007] Furthermore, step S2 includes the following sub-steps: S2.1. Receive test task definition information containing the unique identifier of the sample and target virtual machine identifier information, and parse the operation instruction content and target resource information in the test task definition information; S2.2. Generate a unique identifier for the test task, and query the running status of the corresponding virtual machine based on the target virtual machine identifier information to distinguish between the running status and the stopped status of the virtual machine; S2.3. Perform a startup operation on the virtual machine that is in a stopped state, and repeatedly check the network connection status of the virtual machine at set time intervals until the network connection is successful or the set number of attempts is reached; S2.4. Mark the virtual machine detected by network connection as ready, associate the virtual machine network address information with the unique identifier of the test task, and update the status information of the corresponding virtual machine in the database storage medium.

[0008] Furthermore, step S3 includes the following sub-steps: S3.1. Generate an automated orchestration script based on the test task definition information, map each operation instruction to the corresponding orchestration task node, and arrange the orchestration task nodes in the order of the operation sequence; S3.2. Transfer the malicious application sample file and the automated orchestration script to the specified storage path of the target virtual machine via a proxyless network connection; S3.3. Identify the platform type information corresponding to the target virtual machine, and match the corresponding operation execution logic according to the platform type information; S3.4. Execute the corresponding test operations sequentially according to the operation sequence in the automated orchestration script, and send the execution result data of each operation back to the data storage terminal in real time.

[0009] Furthermore, step S4 includes the following sub-steps: S4.1. Start a traffic capture process before the malicious application sample file runs to monitor the network interface data inside the target virtual machine; S4.2. During the test operation, continuously collect network data packets, generate raw traffic data files, and record the time range information corresponding to the traffic collection; S4.3. After the test operation is completed, stop the traffic capture process and reclaim the original traffic data file and test execution log data to the backend storage location; S4.4. Parse the raw traffic data file to extract network behavior features, including communication connection features and traffic statistics features, and integrate the network behavior features to generate structured analysis data.

[0010] Furthermore, in step S2, a corresponding subtask is generated for each target virtual machine based on the target virtual machine selection information and the unique identifier of the test task. Corresponding test resources are allocated to the subtask, the execution status of the subtask is maintained independently, and the test process corresponding to all subtasks is executed in parallel.

[0011] Furthermore, in step S3, corresponding underlying operation technologies are pre-encapsulated for different operating system distributions and different application types. The underlying operation technologies include command-line interaction technology and graphical interface event injection technology, providing a unified calling interface. The corresponding underlying operation technology is called to execute operation instructions according to the platform type information of the target virtual machine.

[0012] Furthermore, in step S4, during the process of parsing network traffic data, communication connection information is extracted from the network traffic data. The communication connection information includes source address information, destination address information, port information, and protocol type information. Traffic statistical feature information is also collected, including the total number of data packets and the total number of data bytes. The communication connection information and traffic statistical feature information are then integrated to generate structured analysis data.

[0013] Furthermore, in step S4, after generating structured analysis data, the structured analysis data and test execution result data are pushed to the front-end interactive interface through real-time communication. The network connection information and test execution log information are displayed in the form of visual charts, and the set abnormal network connection information is displayed.

[0014] An automated testing system for malicious applications on ARM servers is provided. This system includes a front-end interaction module, a back-end service module, a virtualization resource management module, an automated task execution module, a multi-platform operation encapsulation module, and a network traffic analysis module. The front-end interaction module receives user input and displays data processing results. The back-end service module handles business logic requests and stores system metadata. The virtualization resource management module manages the lifecycle and running status of ARM architecture virtual machines. The automated task execution module generates automated orchestration scripts and distributes test tasks for execution. The multi-platform operation encapsulation module adapts to the underlying operating technologies of different platforms and provides a unified calling interface. The network traffic analysis module captures and parses network traffic data and generates structured analysis data.

[0015] The beneficial effects of this invention are: (1) By constructing a virtualized test environment adapted to the ARM architecture, the complete process of sample management, task scheduling, execution control and traffic analysis is connected to realize the automated and orderly conduct of dynamic testing of malicious applications in the ARM server scenario; (2) Unify the underlying operation technology of different operating environments, provide standardized calling interfaces, eliminate the adaptation differences of different operating system distributions, and improve the reusability of cross-platform test scripts and system usability; (3) Capture and structured parsing of network traffic inside the virtual machine, covering internal communication scenarios that cannot be obtained by external packet capture methods, enriching the analysis dimensions of malicious application network behavior, and improving the comprehensiveness of behavior analysis. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the steps of an automated testing method for malicious applications targeting ARM servers; Figure 2 The following is a flowchart illustrating the specific steps of an automated testing method for malicious applications targeting ARM servers, provided as an example. Detailed Implementation

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1 See Figure 1 This embodiment provides an automated testing method for malicious applications targeting ARM servers, which includes the following steps: S1. Receive malicious application sample files and platform type information, calculate the hash value of the malicious application sample files to generate a unique sample identifier, store the malicious application sample files, and record the metadata information corresponding to the malicious application sample files. S2. Receive test task definition information and target virtual machine selection information containing the unique identifier of the sample, generate a unique identifier of the test task, confirm the running status of the virtual machine, perform a startup operation on virtual machines that are not running, detect the network connection status of the virtual machine, and associate the network address information of the virtual machine that passes the detection with the unique identifier of the test task. S3. Generate an automated orchestration script based on the test task definition information, distribute the automated orchestration script and malicious application sample file to the target virtual machine, match the corresponding operation execution logic, and execute the operation sequence corresponding to the automated orchestration script; S4. Capture network traffic data inside the target virtual machine during the execution of the malicious application sample file, collect test execution result data and network traffic data, parse the network traffic data to generate structured analysis data, and store the structured analysis data and test execution result data.

[0019] In some embodiments, step S1 includes the following sub-steps: S1.1. Receive malicious application sample files and platform type selection information, parse the file data and form data in the upload request, and extract the binary content of the malicious application sample files and platform type selection information; S1.2. Perform a hash operation on the binary content of the malicious application sample file to generate a unique sample identifier for unique identification of the sample; S1.3. Store the malicious application sample files to the corresponding storage locations according to the set path rules, and record the storage path information corresponding to the malicious application sample files; S1.4. Summarize the metadata information of malicious application sample files. The metadata information includes the sample's unique identifier, file size information, upload time information, platform type information, and storage path information. Write the metadata information into the database storage medium.

[0020] In some embodiments, step S2 includes the following sub-steps: S2.1. Receive test task definition information containing the unique identifier of the sample and target virtual machine identifier information, and parse the operation instruction content and target resource information in the test task definition information; S2.2. Generate a unique identifier for the test task, and query the running status of the corresponding virtual machine based on the target virtual machine identifier information to distinguish between the running status and the stopped status of the virtual machine; S2.3. Perform a startup operation on the virtual machine that is in a stopped state, and repeatedly check the network connection status of the virtual machine at set time intervals until the network connection is successful or the set number of attempts is reached; S2.4. Mark the virtual machine detected by network connection as ready, associate the virtual machine network address information with the unique identifier of the test task, and update the status information of the corresponding virtual machine in the database storage medium.

[0021] In some embodiments, step S3 includes the following sub-steps: S3.1. Generate an automated orchestration script based on the test task definition information, map each operation instruction to the corresponding orchestration task node, and arrange the orchestration task nodes in the order of the operation sequence; S3.2. Transfer the malicious application sample file and the automated orchestration script to the specified storage path of the target virtual machine via a proxyless network connection; S3.3. Identify the platform type information corresponding to the target virtual machine, and match the corresponding operation execution logic according to the platform type information; S3.4. Execute the corresponding test operations sequentially according to the operation sequence in the automated orchestration script, and send the execution result data of each operation back to the data storage terminal in real time.

[0022] In some embodiments, step S4 includes the following sub-steps: S4.1. Start a traffic capture process before the malicious application sample file runs to monitor the network interface data inside the target virtual machine; S4.2. During the test operation, continuously collect network data packets, generate raw traffic data files, and record the time range information corresponding to the traffic collection; S4.3. After the test operation is completed, stop the traffic capture process and reclaim the original traffic data file and test execution log data to the backend storage location; S4.4. Parse the raw traffic data file to extract network behavior features, including communication connection features and traffic statistics features, and integrate the network behavior features to generate structured analysis data.

[0023] In some embodiments, in step S2, a corresponding subtask is generated for each target virtual machine based on the target virtual machine selection information and the unique identifier of the test task, corresponding test resources are allocated to the subtask, the execution status of the subtask is maintained independently, and the test processes corresponding to all subtasks are executed in parallel.

[0024] In some embodiments, in step S3, corresponding underlying operation technologies are pre-encapsulated for different operating system distributions and different application types. The underlying operation technologies include command-line interaction technology and graphical interface event injection technology, providing a unified calling interface, and calling the corresponding underlying operation technology to execute operation instructions according to the platform type information of the target virtual machine.

[0025] In some embodiments, during step S4, in the process of parsing network traffic data, communication connection information is extracted from the network traffic data. The communication connection information includes source address information, destination address information, port information, and protocol type information. Traffic statistical feature information is also collected, including the total number of data packets and the total number of data bytes. The communication connection information and traffic statistical feature information are then integrated to generate structured analysis data.

[0026] In some embodiments, after generating structured analysis data in step S4, the structured analysis data and test execution result data are pushed to the front-end interactive interface through real-time communication. The network connection information and test execution log information are displayed in the form of visual charts, and the set abnormal network connection information is displayed.

[0027] In some embodiments, an automated testing system for malicious applications on ARM servers is provided. This system includes a front-end interaction module, a back-end service module, a virtualization resource management module, an automated task execution module, a multi-platform operation encapsulation module, and a network traffic analysis module. The front-end interaction module receives user input and displays data processing results. The back-end service module processes business logic requests and stores system metadata. The virtualization resource management module manages the lifecycle and running status of ARM architecture virtual machines. The automated task execution module generates automated orchestration scripts and distributes test tasks for execution. The multi-platform operation encapsulation module adapts to the underlying operating technologies of different platforms and provides a unified calling interface. The network traffic analysis module captures and parses network traffic data and generates structured analysis data.

[0028] Example 2 This embodiment provides a specific implementation process for an automated testing method for malicious applications on ARM servers. By constructing a virtualized testing environment adapted to the ARM architecture, and combining automated orchestration mechanisms with multi-platform operation adaptability, it achieves automated, large-scale dynamic testing and network behavior analysis of malicious application samples, providing support for malicious application security analysis in ARM server scenarios. Figure 2 As shown, the specific implementation process is as follows: Step 1. Sample Upload and Metadata Storage: Step 1.1. Sample Data Reception and Parsing: The malicious application sample file is an executable program file adapted to the ARM instruction set architecture, capable of running on ARM architecture server operating systems. Platform type information identifies the operating system distribution category to which the sample is adapted, providing a basis for subsequent platform adaptation. The upload request uses a form submission method for data transmission, containing the binary data of the malicious application sample file and platform type selection information. The upload request is parsed, separating the file data and form field content from the request, extracting the binary content of the malicious application sample file and the platform type selection information, providing a data foundation for subsequent identification generation and storage operations.

[0029] Step 1.2. Sample Unique Identifier Generation: Hash operation is a one-way hashing technique in cryptography, a mature and readily available technology. It can convert binary data of arbitrary length into a fixed-length feature string, and the probability of different input data generating the same output result is extremely low. It can be used for unique identification and integrity verification of data. A hash operation is performed on the binary content of the malicious application sample file to generate a unique sample identifier. This unique sample identifier can be applied to multiple stages, including sample deduplication, storage path mapping, and task association, ensuring that each sample has a unique identification basis in the system.

[0030] Step 1.3. Sample File Storage and Path Recording: Store malicious application sample files to their corresponding storage locations according to the set path rules. The path rules use the prefix characters of the sample's unique identifier as the basis for directory partitioning, distributing sample files across different levels of directories to avoid storage management issues caused by an excessive number of files in a single directory. Record the storage path information corresponding to the malicious application sample files; the storage path information corresponds one-to-one with the sample's unique identifier, for use in retrieving sample files during subsequent test task execution.

[0031] Step 1.4. Metadata Aggregation and Persistent Storage: Aggregate the metadata information of the malicious application sample files. This metadata includes the sample's unique identifier, file size, upload time, platform type, and storage path. The database storage medium is a non-relational document database, a mature existing data storage technology that supports flexible storage and fast querying of semi-structured data and can adapt to the system's metadata field expansion needs. The metadata information is written to the database storage medium, completing the sample information registration process. All subsequent operations related to the sample can retrieve the corresponding metadata content by querying the sample's unique identifier.

[0032] In some specific implementations, three types of verification rules are set up for the pre-entry verification of malicious application sample files to screen uploaded sample files layer by layer, preventing invalid files, corrupted files, or files with non-target architectures from entering the sample library and consuming storage and computing resources. The three types of verification rules are file format verification, file size threshold verification, and file architecture matching verification. File format verification identifies the executable format type of the file by reading the fixed identifier field in the file header, confirming that the file belongs to the standard ELF executable file format. Files that do not meet the format requirements directly terminate the upload process and return the corresponding prompt.

[0033] File size threshold verification sets upper and lower limits for uploaded file sizes. Empty or invalid files below the lower limit, and excessively large files exceeding the upper limit, are not accepted, reducing the burden on the storage system and subsequent analysis processes caused by extreme-sized files. File architecture matching verification verifies the instruction set architecture of the file by parsing the architecture identifier field in the executable file header. Sample files with non-ARM architecture cannot run in the test environment, so they are directly intercepted and the reason for the architecture mismatch is marked. The results of all verification steps generate verification logs, which are stored in association with the sample upload request information, facilitating subsequent traceability of the sample entry verification process. This implementation method can complete the initial quality screening at the sample entry stage, increase the proportion of valid samples in the sample library, reduce the probability of subsequent test tasks failing due to invalid samples, and ensure the stability of the overall test process.

[0034] In some specific implementations, a two-level directory storage architecture is adopted for the storage and indexing management of malicious application sample files. The first-level directory is named using the first two characters of the sample's unique identifier, and the second-level directory is named using the complete sample's unique identifier. Each malicious application sample file corresponds to an independent second-level storage directory. The sample's unique identifier is generated using a combination of two hash algorithms: MD5 hash value and SHA256 hash value. The SHA256 hash value serves as the primary unique identifier of the sample, and the MD5 hash value serves as an auxiliary deduplication identifier. The metadata fields written to the database storage medium contain six types of content: the primary unique identifier of the sample, the auxiliary unique identifier of the sample, the sample file size information, the sample upload time information, the sample's compatible platform type information, and the sample file storage path information. A deduplication check is performed on each uploaded sample file. By comparing the primary unique identifier of the sample already stored in the database, it is confirmed whether the currently uploaded sample already exists in the sample library. If it already exists, the file storage process is skipped, and only the sample's association record information is updated. This implementation method can manage massive sample files in an orderly manner, avoiding the problem of decreased access efficiency caused by too many files in a single directory. At the same time, it reduces the probability of sample collisions by combining double hash identifiers, thereby improving the stability and reliability of sample library management.

[0035] In some embodiments, an integrity verification step can be added before storing the sample file. A verification value is generated by a secondary hash operation and compared with the verification value at the upload end to confirm that the file has not been damaged during transmission. The storage operation is then performed after the verification is passed, thus ensuring the integrity of the stored sample file.

[0036] Step 2. Task Definition and Virtual Machine Resource Preparation: Step 2.1. Task Information Reception and Parsing: Receive test task definition information and target virtual machine identifier information, both containing unique sample identifiers. The test task definition information contains an ordered set of operation instructions, specifying the complete action flow for test execution. The target virtual machine identifier information specifies the virtualization resources on which the test task will run. Parse the test task definition information and target virtual machine identifier information to extract the operation instruction content and target resource information from the test task definition information. This clarifies the action sequence for test execution and the corresponding target virtual machine resources, providing a basis for subsequent task orchestration and resource scheduling.

[0037] Step 2.2. Task Identifier Generation and Status Query: Generate a unique identifier for each test task to distinguish between different test tasks and enable status tracking and data association throughout the task's lifecycle. Query the running status of the corresponding virtual machine based on the target virtual machine identifier information, distinguishing between running and stopped states to provide a basis for subsequent environment preparation operations. A running virtual machine status indicates that the virtualization instance is powered on and ready for task deployment; a stopped virtual machine status indicates that the virtualization instance is powered off and requires a startup operation first.

[0038] Step 2.3. Virtual Machine Startup and Connection Detection: Perform a startup operation on the virtual machine in the stopped state. The virtualization environment implements the operation of ARM architecture virtual machines based on hardware-assisted virtualization technology, and manages the lifecycle of the virtual machine uniformly through the libvirt interface. libvirt is an existing general-purpose virtualization management interface that can adapt to various underlying virtualization technologies and provide unified virtual machine control and status query capabilities. In this embodiment, the startup, shutdown, and status query operations of the virtual machine are performed through the libvirt interface to achieve standardized management of virtualization resources. The network connection status of the virtual machine is repeatedly checked at set time intervals until the network connection is successful or the set number of attempts is reached, confirming that the network service inside the virtual machine is in an available state, providing a network connectivity foundation for subsequent task deployment.

[0039] Step 2.4. Ready Status Marking and Resource Association: Mark the virtual machine detected through network connection as ready, associate the virtual machine's network address information with the unique identifier of the test task, and establish a mapping relationship between the test task and the execution resource. Update the status information of the corresponding virtual machine in the database storage medium to synchronize the resource usage status and avoid conflicts caused by multiple tasks occupying the same virtualization resource simultaneously.

[0040] After associating virtual machine resources with test tasks, a corresponding subtask is generated for each target virtual machine based on the target virtual machine selection information and the unique identifier of the test task. Corresponding test resources are allocated to each subtask, and the execution status of each subtask is maintained independently. All test processes corresponding to the subtasks are executed in parallel. By executing multiple subtasks in parallel, the overall processing efficiency of the test tasks is improved, adapting to the needs of batch sample and multi-environment parallel testing.

[0041] In some specific implementations, to address the need for rapid reset and reuse of virtualized test environments, three virtual machine snapshot management strategies are configured to enable flexible saving and rollback of the test environment state, reducing the workload of repeatedly building test environments. The three snapshot management strategies are baseline template snapshot, task pre-installation snapshot, and cleanup completion snapshot. The baseline template snapshot is generated based on a pre-configured ARM server operating system template, preserving the clean state after the initial operating system installation. All newly created test virtual machines are quickly generated by cloning this baseline template snapshot, eliminating the need to perform operating system installation and basic configuration operations on each machine.

[0042] A pre-task snapshot is generated before the test task deploys sample files, recording the system state before the test environment deployment. If a system anomaly or sample corruption occurs during test task execution, the system can be rolled back to the state corresponding to the pre-task snapshot, quickly restoring a usable test environment without recreating virtual machine instances. A cleanup completion snapshot is generated after the test task ends and the environment cleanup operation is completed. It is used to verify the effectiveness of the environment cleanup and can also serve as the starting state for subsequent test tasks of the same type, shortening the test environment preparation cycle. All snapshot generation, rollback, and deletion operations are executed through the virtualization management interface. Snapshot information is associated with the corresponding virtual machine identifier and stored, recording the snapshot generation time and corresponding state description. This implementation significantly shortens the preparation and reset time of the test environment, improves the reuse efficiency of virtualization resources, and ensures that each test task can be executed in a consistent environment, improving the reproducibility of test results.

[0043] In some specific implementations, a fixed detection cycle and retry limit are set for the network connectivity detection step after virtual machine startup to ensure the stability of the test environment readiness determination. The detection interval is set to 3 seconds, and a network connection attempt is performed once per detection. The maximum number of retries is set to 20. Virtual machines that successfully establish a network connection within the retry limit are marked as ready; virtual machines that fail to establish a network connection after exceeding the maximum number of retries are marked as having a connection timeout, and the corresponding subtask is terminated and the exception information is recorded. Virtual machine states are divided into four categories: stopped, starting, ready, and abnormal. Each state corresponds to different task operation permissions. Virtual machines in the stopped state can only perform startup operations, virtual machines in the starting state can only perform status query operations, virtual machines in the ready state can receive test task deployments, and virtual machines in the abnormal state require manual investigation and state reset. This implementation method can verify the service readiness of virtual machines with a fixed detection rhythm, avoiding task blocking caused by unlimited waiting. At the same time, through the fine-grained division of multiple states, it can achieve full lifecycle management of virtualization resources and improve the accuracy and orderliness of test task scheduling.

[0044] In some embodiments, the allocation order of subtasks can be dynamically adjusted according to the resource load of the virtual machine, and test tasks can be allocated to virtual machine nodes in an idle state first, so as to balance the resource load of each node and improve the overall utilization efficiency of virtualization resources.

[0045] Step 3. Task orchestration and automated execution: Step 3.1. Automated Orchestration Script Generation: An automated orchestration script is generated based on the test task definition information. This script uses a declarative configuration format to describe the task execution flow, offering strong readability and cross-platform adaptability. Each operation instruction is mapped to a corresponding orchestration task node, which is arranged in the order of the operation sequence to ensure that test actions are executed sequentially according to the preset logical order. In this embodiment, automated orchestration capabilities are implemented using the Ansible tool. Ansible is a mature automated operation and maintenance orchestration tool that adopts an agentless architecture and uses the SSH protocol to achieve remote task execution and management, adapting to the task distribution needs of batch nodes.

[0046] Step 3.2. Test Resource Distribution and Transmission: The malicious application sample file and automated orchestration script are transmitted to a specified storage path on the target virtual machine via an agentless network connection. The agentless network connection method eliminates the need for a resident agent program to be pre-deployed within the target virtual machine; data transmission and command execution can be completed solely through the system's native network services. This reduces intrusion into the test environment, minimizes interference from additional programs on the malicious application's behavior, and ensures the objectivity of the test results.

[0047] Step 3.3. Platform Type Identification and Execution Logic Matching: Identify the platform type information corresponding to the target virtual machine, and match the corresponding operation execution logic based on the platform type information. In this embodiment, corresponding underlying operation technologies are pre-encapsulated for different operating system distributions and different application types. The underlying operation technologies include command-line interaction technology and graphical interface event injection technology, providing a unified calling interface upwards.

[0048] Command-line interaction technology is an existing technique that simulates command-line input and output interaction through scripts, used to adapt to server operating environments without a graphical interface. Graphical interface event injection technology is an existing technique that injects input events into a graphical interface system, used to adapt to application operating scenarios with a graphical interface. Based on the platform type information of the target virtual machine, the corresponding underlying operation technology is called to execute operation instructions. Upper-layer tasks do not need to concern themselves with the implementation differences of the underlying platform; they only need to call standardized interfaces to complete cross-platform operation execution, improving the reusability of test scripts.

[0049] Step 3.4. Operation Sequence Execution and Result Feedback: Execute the corresponding test operations sequentially according to the operation sequence in the automated orchestration script, and feed back the execution result data of each operation to the data storage terminal in real time. The execution result data includes the execution status of the operation, standard output content, and error output content. Providing feedback on the execution status and output information step by step facilitates tracking the test execution progress and troubleshooting any anomalies that occur during execution.

[0050] In some specific implementations, for task scheduling scenarios of large-scale parallel testing, the automated task execution engine adopts a multi-threaded task distribution architecture and sets an upper limit parameter for the number of concurrent tasks. A single control node supports a maximum of 32 parallel subtasks simultaneously. Subtasks exceeding this limit enter a task waiting queue and are scheduled for execution sequentially according to their submission order. The operation instructions for test tasks are divided into five categories: file transfer instructions, command execution instructions, traffic control instructions, delay instructions, and environment cleanup instructions. Each category corresponds to an independent encapsulation interface and execution logic. When generating the automated orchestration script, different types of instructions are mapped to corresponding orchestration task nodes according to the operation sequence. A step-by-step status callback mechanism is used during task execution. After each operation instruction is executed, the execution status and output content are sent back to the data storage terminal. The callback content includes three types of information: step number, execution status code, and output text content. This implementation method can schedule batch test tasks in an orderly manner, avoiding the problem of control node resource overload caused by excessive concurrency. At the same time, through classified instruction encapsulation and step-by-step callback mechanism, it ensures the controllability and traceability of task execution, and adapts to the use requirements of multi-node parallel testing.

[0051] In some embodiments, an exception retry mechanism can be added during the operation execution process. When a single-step operation fails, the corresponding operation is re-executed according to the set rules, which reduces the impact of temporary network fluctuations or environmental anomalies on the test task and improves the success rate of the test task execution.

[0052] Step 4. Traffic Collection and Result Analysis and Storage: Step 4.1. Traffic Capture Process Startup: The traffic capture process is started before the malicious application sample file runs to monitor the network interface data inside the target virtual machine. This embodiment uses the Scapy tool to implement the traffic capture function. Scapy is a mature network packet processing tool that supports packet capture, construction, parsing, and statistical operations, and can flexibly adapt to customized traffic collection needs. The traffic capture process runs inside the target virtual machine and can directly obtain traffic data from all network interfaces inside the virtual machine, including communication data from the loopback interface and virtual network interfaces, covering internal communication scenarios that cannot be captured by external bypass packet capture methods.

[0053] In some specific implementations, four types of configurable collection parameters are predefined for the initial configuration phase of the traffic capture process, enabling flexible adaptation of collection strategies under different test scenarios. These four types of collection parameters are network interface selection parameters, packet filtering rules, single file fragmentation thresholds, and process protection timeout parameters. The network interface selection parameter supports specifying three collection ranges: single physical network card collection, all virtual interface collection, and full network card collection. The default is the full network card collection mode, covering the physical network cards inside the virtual machine, loopback interfaces, and container virtual network interfaces, avoiding the omission of communication data from hidden channels. The packet filtering rules have three preset templates: a full unfiltered template, a template excluding internal network addresses, and a TCP / UDP protocol-only template. When creating a test task, the corresponding template can be selected according to the analysis target. The filtering rules are implemented based on standard packet filtering syntax, directly discarding packets that do not conform to the rules during the capture phase, reducing the storage and transmission overhead of invalid data. The single-file fragment threshold controls the maximum size of a single raw traffic file. When the file size reaches the threshold during collection, the current file is automatically closed and a new traffic file is generated. All fragment files are associated with the same test task identifier to avoid decreased reading and transmission efficiency due to excessively large single file sizes. The process guardian timeout parameter sets the maximum runtime of the traffic capture process. When the test task encounters an anomaly and cannot issue a stop command normally, the capture process automatically terminates after reaching the timeout threshold to prevent the process from resident for a long time and occupying virtual machine computing resources.

[0054] All acquisition parameters are transmitted synchronously along with the test task definition information and are loaded and take effect when the traffic capture process starts. Parameter configuration records are stored in association with the unique identifier of the test task. This implementation method can adjust the scope and granularity of traffic acquisition according to different test requirements. While ensuring the complete acquisition of critical communication data, it reduces the consumption of storage and analysis resources by redundant data. At the same time, it avoids resource leakage in abnormal scenarios through a timeout protection mechanism, thereby improving the stability and adaptability of the traffic acquisition process.

[0055] Step 4.2. Continuous Traffic Data Collection: During the test operation, network data packets are continuously collected to generate raw traffic data files, recording the time range information corresponding to the traffic collection. The raw traffic data files retain the complete binary content of the data packets, without filtering or deletion, and can support subsequent multi-dimensional in-depth analysis and retrospective verification.

[0056] In some specific implementations, a two-level caching mechanism is used to handle the temporary data management stage during traffic collection, balancing collection performance and data integrity. The two levels of caching are a memory caching layer and a local disk caching layer. The memory caching layer has a fixed queue length, with a single batch of 128 data packets cached. Collected data packets are first written to the memory caching queue. When the queue length reaches the threshold, they are written in batches to the original traffic files in the local disk caching layer. This reduces the impact of frequent disk I / O on virtual machine performance and minimizes interference from the collection process on the runtime environment of malicious applications.

[0057] The local disk caching layer generates raw traffic data files using an append-only write method. During file writing, timestamps of data packets are recorded synchronously to ensure accurate time order of traffic data, with timestamp precision consistent with the system clock. A snapshot of the collection status is generated at fixed intervals during the collection process. The snapshot content includes the number of collected data packets, the number of bytes written, and the current running status. These status snapshots are transmitted back to the backend in real time via a network connection for monitoring the collection process. If the collection process exits abnormally, the range of collected data can be confirmed based on the last status snapshot. The generation interval for the collection status snapshot is set to 10 seconds. Snapshot data is associated with a unique identifier for the corresponding test task for subsequent task status tracing and anomaly investigation. This implementation reduces the performance impact of traffic collection on the test environment through a layered caching mechanism, ensures the objectivity of the malicious application's operating environment, and enables monitorability of the collection process through periodic status snapshots, facilitating timely detection and handling of collection anomalies and improving the reliability of the traffic collection process.

[0058] Step 4.3. Data Acquisition Stop and Result Reclamation: After the test operation is completed, stop the traffic capture process and reclaim the original traffic data files and test execution log data to the backend storage location. The test execution log data contains the execution output and status records of all operations during the test, which can completely reconstruct the test execution process and provide auxiliary basis for subsequent behavior analysis. After the data reclamation is completed, the temporary sample files and test data in the target virtual machine can be cleaned up to restore the initial state of the test environment and provide a clean running environment for subsequent test tasks.

[0059] Step 4.4. Traffic Data Parsing and Structured Processing: Parse the raw traffic data file to extract network behavior features, including communication connection features and traffic statistics features. Integrate these network behavior features to generate structured analysis data. During the parsing of network traffic data, communication connection information is extracted, including source address information, destination address information, port information, and protocol type information. Traffic statistics features are also extracted, including the total number of data packets and the total number of data bytes. Integrate the communication connection information and traffic statistics features to generate structured analysis data. The structured analysis data is stored using a unified field format for easy subsequent querying, statistics, and visualization.

[0060] After generating structured analysis data, this data, along with test execution results, is pushed to the front-end interactive interface via real-time communication. Network connection information and test execution logs are displayed in visual charts, along with any abnormal network connection information. Real-time communication supports instant data push, allowing users to simultaneously view test progress and analysis results without waiting for the entire task to complete. The front-end interactive interface employs a front-end / back-end separation architecture, obtaining back-end data through standardized interfaces to implement test task management, sample management, and analysis result display functions.

[0061] In some specific implementations, to address the full lifecycle storage management needs of network traffic data, a two-tier storage architecture is set up to classify and store traffic data with different usage frequencies. Simultaneously, four types of data access permissions are set for security control, balancing data access efficiency, storage costs, and data security. The two-tier storage architecture consists of an online hot storage layer and an offline cold storage layer. The online hot storage layer stores test task data within a defined time range, including structured analysis data, index information of raw traffic files, and test execution logs. It supports fast querying and visualization, meeting the high-frequency access needs of daily malicious application analysis. The offline cold storage layer stores historical data exceeding the retention period of hot storage, including raw traffic data files and historical analysis reports, for long-term archiving and retrospective analysis. Data access to the cold storage layer requires a prior retrieval request, and normal access is only possible after the data is migrated to the hot storage layer by the system.

[0062] Data migration between the two-tier storage is performed automatically according to a set cycle. During the migration process, the complete association identifiers and metadata information of the data are preserved to ensure that the retrieved data can be accurately matched with the corresponding test tasks and sample information. Four types of data access permissions are provided: read-only viewing permission, result export permission, sample management permission, and system configuration permission. Read-only viewing permission only allows browsing of analysis results and visualization content; result export permission allows downloading analysis reports and raw traffic files; sample management permission allows performing operations such as uploading and deleting samples; and system configuration permission allows adjusting storage policies and permission rules. This implementation method enables differentiated storage management of data with different usage frequencies. While ensuring the efficiency of daily analysis and access, it reduces the resource investment in long-term data storage. Simultaneously, hierarchical permission control reduces the risk of data leakage and misoperation, improving the security and rationality of traffic data management.

[0063] In some specific implementations, regarding the parsing and processing of network traffic data, seven types of structured network features are extracted after parsing the raw traffic data file, covering multiple dimensions of network communication. These seven network features are: source network address information, destination network address information, transport layer port information, transport protocol type information, packet quantity information, total data byte count information, and application layer protocol identifier information. The application layer protocol identifier information is identified through both port matching and payload feature matching, enabling the identification of common network service protocols and malicious communication protocol types.

[0064] The parsing process follows a layered processing logic. First, it parses the data link layer and network layer headers to extract addresses and protocol content. Next, it parses the transport layer header to extract port numbers and message lengths. Finally, it extracts payload fragments for application layer feature matching. The generated structured analysis data is stored in a database as documents, with each document associated with a unique identifier for the corresponding test task and a target virtual machine identifier. This implementation extracts multi-dimensional structured features from raw traffic data, converting unstructured binary traffic data into queryable and statistically standardized data. This facilitates subsequent behavior analysis and visualization, improving the efficiency and ease of use of network behavior analysis.

[0065] In some embodiments, an application layer protocol deep parsing function can be added during the traffic parsing process to extract key field content of the application layer protocol, further enrich the analysis dimensions of malicious application network behavior, and support more detailed behavior judgment work.

[0066] This embodiment constructs a virtualized testing environment adapted to the ARM architecture, providing a native runtime testing environment for ARM-based malicious applications. This alleviates the compatibility limitations of traditional x86 testing platforms that cannot directly run ARM malicious code, ensuring the effectiveness of dynamic behavior analysis. By unifying the operation logic across multiple platforms, it eliminates operational differences between different operating system distributions, reduces the development and maintenance costs of cross-platform test scripts, and improves the reusability of the test solution and the ease of use of the system. Relying on an automated orchestration mechanism, it achieves batch distribution and parallel execution of test tasks, improving the overall execution efficiency of malicious application testing and adapting to the needs of large-scale sample analysis. It adopts a virtual machine internal traffic collection method, covering internal network communication scenarios that traditional external packet capture methods cannot obtain, improving the analysis dimensions of malicious application network behavior and enhancing the comprehensiveness of behavior analysis. The overall solution adopts a modular design approach, with each functional module interacting through standardized interfaces. It has a certain degree of scalability, allowing for the addition of adapted platform types and analysis functions according to actual needs, adapting to constantly changing malicious application analysis scenarios. By combining and implementing the above technical solutions, automated and large-scale testing support can be provided for malicious applications on ARM servers, improving the efficiency and comprehensiveness of malicious application behavior analysis and providing data support for security protection in ARM server scenarios.

[0067] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. An automated testing method for malicious applications targeting ARM servers, characterized in that, The method includes the following steps: S1. Receive malicious application sample files and platform type information, calculate the hash value of the malicious application sample files to generate a unique sample identifier, store the malicious application sample files, and record the metadata information corresponding to the malicious application sample files. S2. Receive test task definition information and target virtual machine selection information containing the unique identifier of the sample, generate a unique identifier of the test task, confirm the running status of the virtual machine, perform a startup operation on virtual machines that are not running, detect the network connection status of the virtual machine, and associate the network address information of the virtual machine that passes the detection with the unique identifier of the test task. S3. Generate an automated orchestration script based on the test task definition information, distribute the automated orchestration script and malicious application sample file to the target virtual machine, match the corresponding operation execution logic, and execute the operation sequence corresponding to the automated orchestration script; S4. Capture network traffic data inside the target virtual machine during the execution of the malicious application sample file, collect test execution result data and network traffic data, parse the network traffic data to generate structured analysis data, and store the structured analysis data and test execution result data.

2. The method according to claim 1, characterized in that, Step S1 includes the following sub-steps: S1.

1. Receive malicious application sample files and platform type selection information, parse the file data and form data in the upload request, and extract the binary content of the malicious application sample files and platform type selection information; S1.

2. Perform a hash operation on the binary content of the malicious application sample file to generate a unique sample identifier for unique identification of the sample; S1.

3. Store the malicious application sample files to the corresponding storage locations according to the set path rules, and record the storage path information corresponding to the malicious application sample files; S1.

4. Summarize the metadata information of malicious application sample files. The metadata information includes the sample's unique identifier, file size information, upload time information, platform type information, and storage path information. Write the metadata information into the database storage medium.

3. The method according to claim 1, characterized in that, Step S2 includes the following sub-steps: S2.

1. Receive test task definition information containing the unique identifier of the sample and target virtual machine identifier information, and parse the operation instruction content and target resource information in the test task definition information; S2.

2. Generate a unique identifier for the test task, and query the running status of the corresponding virtual machine based on the target virtual machine identifier information to distinguish between the running status and the stopped status of the virtual machine; S2.

3. Perform a startup operation on the virtual machine that is in a stopped state, and repeatedly check the network connection status of the virtual machine at set time intervals until the network connection is successful or the set number of attempts is reached; S2.

4. Mark the virtual machine detected by network connection as ready, associate the virtual machine network address information with the unique identifier of the test task, and update the status information of the corresponding virtual machine in the database storage medium.

4. The method according to claim 1, characterized in that, Step S3 includes the following sub-steps: S3.

1. Generate an automated orchestration script based on the test task definition information, map each operation instruction to the corresponding orchestration task node, and arrange the orchestration task nodes in the order of the operation sequence; S3.

2. Transfer the malicious application sample file and the automated orchestration script to the specified storage path of the target virtual machine via a proxyless network connection; S3.

3. Identify the platform type information corresponding to the target virtual machine, and match the corresponding operation execution logic according to the platform type information; S3.

4. Execute the corresponding test operations sequentially according to the operation sequence in the automated orchestration script, and send the execution result data of each operation back to the data storage terminal in real time.

5. The method according to claim 1, characterized in that, Step S4 includes the following sub-steps: S4.

1. Start a traffic capture process before the malicious application sample file runs to monitor the network interface data inside the target virtual machine; S4.

2. During the test operation, continuously collect network data packets, generate raw traffic data files, and record the time range information corresponding to the traffic collection; S4.

3. After the test operation is completed, stop the traffic capture process and reclaim the original traffic data file and test execution log data to the backend storage location; S4.

4. Parse the raw traffic data file to extract network behavior features, including communication connection features and traffic statistics features, and integrate the network behavior features to generate structured analysis data.

6. The method according to claim 1, characterized in that, In step S2, a corresponding subtask is generated for each target virtual machine based on the target virtual machine selection information and the unique identifier of the test task. Corresponding test resources are allocated to the subtask, the execution status of the subtask is maintained independently, and the test process corresponding to all subtasks is executed in parallel.

7. The method according to claim 1, characterized in that, In step S3, corresponding underlying operation technologies are pre-encapsulated for different operating system distributions and different application types. The underlying operation technologies include command-line interaction technology and graphical interface event injection technology, providing a unified calling interface. The corresponding underlying operation technology is called to execute operation instructions according to the platform type information of the target virtual machine.

8. The method according to claim 1, characterized in that, In step S4, during the parsing of network traffic data, communication connection information is extracted from the network traffic data. The communication connection information includes source address information, destination address information, port information, and protocol type information. Traffic statistical feature information is also collected, including the total number of data packets and the total number of data bytes. The communication connection information and traffic statistical feature information are then integrated to generate structured analysis data.

9. The method according to claim 1, characterized in that, In step S4, after generating structured analysis data, the structured analysis data and test execution result data are pushed to the front-end interactive interface through real-time communication. The network connection information and test execution log information are displayed in the form of visual charts, and the set abnormal network connection information is displayed.

10. An automated testing system for malicious applications targeting ARM servers, used to execute the method as described in any one of claims 1-9, characterized in that, The system includes a front-end interaction module, a back-end service module, a virtualization resource management module, an automated task execution module, a multi-platform operation encapsulation module, and a network traffic analysis module; the front-end interaction module is used to receive user input information and display data processing results. The backend service module is used to process business logic requests and store system metadata; the virtualization resource management module is used to manage the lifecycle and running status of ARM architecture virtual machines; The automated task execution module is used to generate automated orchestration scripts and distribute test tasks for execution; the multi-platform operation encapsulation module is used to adapt to the underlying operation technologies of different platforms and provide a unified calling interface; The network traffic analysis module is used to capture and parse network traffic data and generate structured analysis data.