A Static Scanning Method and System for Virtual Machine Images Based on a Hyper-Converged Platform
By adopting static scanning methods on the hyperconverged platform, including mirror acquisition, resource scheduling and deep scanning of the comprehensive scanning engine, the problem of large-scale static virtual machine image scanning is solved, efficient security risk management is achieved, and the security and stability of the system are ensured.
Patent Information
- Application Number
- CN202510363705.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-26
AI Technical Summary
On the hyper-converged platform, it is difficult for the existing technology to effectively scan and process large-scale static virtual machine images, which makes it difficult to prevent security risks and affect the security, stability and compliance of the system.
A virtual machine mirror static scanning method based on a hyperconverged platform is adopted to create consistent snapshots through the mirror acquisition and preprocessing module, and use index and search algorithms to obtain the image files and extract metadata information, perform multiple rounds of hash calculations and resource scheduling, and conduct multi-stage deep scanning combined with the comprehensive scanning engine to identify potential high-risk vulnerabilities and malware, and conduct comprehensive review of configuration files.
It realizes the completion of virtual machine image scanning in a large-scale hyperconverged cluster in a short time, effectively reducing security risks and ensuring high security, stability and compliance of the virtualized environment.
Smart Images

Figure CN119885168B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of virtual machine image static scanning, and particularly relates to a method and system for static scanning of virtual machine images based on a hyper-converged platform. Background Art
[0002] In the current era of surging digitalization, relying on its excellent resource integration characteristics, the hyper-converged platform, which integrates computing, storage, and network resources, has become one of the core choices for many enterprises to build modern data centers and cloud computing infrastructures. As the fundamental template for constructing diverse virtual machine instances on the hyper-converged platform, the security of virtual machine images undoubtedly plays a crucial role in the entire virtualization ecosystem.
[0003] With the continuous evolution of information technology, the forms of network security threats have become increasingly complex and changeable. Traditional security protection measures for virtual machine images have gradually revealed many insurmountable limitations. In the complex and distributed environment of the hyper-converged platform, existing scanning technologies often perform poorly in terms of scanning completeness, accuracy, rationality of resource utilization, and efficiency in handling large-scale static virtual machine images. These defects are extremely likely to cause severe security threats such as malicious attackers using hidden vulnerabilities in the image for illegal intrusion, malware spreading through the image and then destroying the stability and integrity of the entire system, and network security vulnerabilities caused by improper configuration leading to leakage of sensitive information during the subsequent deployment and operation stages of the virtual machine, bringing inestimable risks and losses to the business operations of enterprises and the protection of data assets.
[0004] Therefore, how to pre-emptively eliminate various potential security hazards that may be hidden before the deployment of virtual machines in the production environment and ensure that the virtualization environment supported by the hyper-converged platform can maintain high security, stability, and compliance is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for static scanning of virtual machine images based on a hyper-converged platform, so as to pre-emptively eliminate various potential security hazards that may be hidden before the deployment of virtual machines in the production environment and ensure that the virtualization environment supported by the hyper-converged platform can maintain high security, stability, and compliance.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A method for static scanning of virtual machine images based on a hyper-converged platform, comprising the following steps:
[0008] S1: Create a consistent snapshot of the mirror storage to be scanned through the mirror acquisition and preprocessing module. Based on the hyper-converged distributed nodes, use the indexing and retrieval algorithm to obtain the static virtual machine mirror file, extract and record its metadata information, perform multiple rounds of hash calculation on the virtual machine mirror file, and compare and verify it with the preset standard hash value;
[0009] S2: Through the resource scheduling and monitoring module, collect the comprehensive resource data of the hyper-converged platform in real time, and use the preset algorithm to allocate corresponding resource quotas for the scanning task, including the optimal execution node of the scanning task, the allocated number of CPU cores, the set memory capacity, and the network bandwidth limit;
[0010] S3: Conduct multi-stage in-depth scanning through the comprehensive scanning engine module. First, conduct preliminary screening with file fingerprints and feature matching, then exclude misjudgments through intelligent version comparison, and finally use static code analysis to deeply dig potential high-risk vulnerabilities for key complex components. Quickly identify and locate known malware by extracting and analyzing multi-dimensional features of various files in the virtual machine mirror, and conduct a comprehensive review of the configuration files in the mirror;
[0011] S4: Through the result processing and report generation module, comprehensively collect various result data generated by the scanning engine module during the scanning process, and use the preset algorithm to deeply classify, organize, correlate, analyze, and summarize and statistics the various result data.
[0012] Preferably, the specific process of creating a consistent snapshot of the mirror storage to be scanned through the mirror acquisition and preprocessing module in step S1 is as follows:
[0013] S11: The mirror acquisition and preprocessing module receives an external start instruction or activates the mirror acquisition and preprocessing module at a specified time node to start executing the task;
[0014] S12: By calling the snapshot creation interface provided by the distributed storage system, generate a consistent snapshot for the target mirror storage when starting the acquisition;
[0015] S13: Set mirror indexing and retrieval based on the hyper-converged distributed nodes, quickly locate and obtain the target mirror in the massive static virtual machine mirrors, obtain the mirror metadata information from the distributed file mirror storage snapshot storage, screen out the specified index items from the snapshot storage to obtain the mirror metadata information, and use the preset hash algorithm to allocate the node ID and the storage location of the index item;
[0016] S14: Create a distributed mirror scanning staging file to store the mirror to be scanned, obtain the list of mirrors to be scanned, record the metadata information of the mirrors to be scanned, judge the format of the mirrors to be scanned, store the mirrors in the newly created distributed mirror scanning staging file storage, and perform mirror integrity calculation.
[0017] Preferably, the specific process of step S13 is as follows:
[0018] S131: Obtain mirror metadata information from the distributed file mirror storage snapshot storage, and screen out key features including the mirror name, creation time, and size as index items;
[0019] S132: Use a preset hash algorithm to allocate node IDs and storage locations for index items. Each node generates a hash value as the node ID based on the comprehensive information of its network address and hardware identifier;
[0020] S133: For the index items, calculate their hash values, and map the index items to the corresponding nodes through the preset hash algorithm;
[0021] S134: When a retrieval request is received, according to the given query conditions, multiple nodes perform parallel retrieval to quickly locate the target mirror index.
[0022] Preferably, the specific process of step S14 is as follows:
[0023] S141: Create a distributed mirror scanning staging file storage to store the mirror to be scanned, and obtain the list of mirrors to be scanned. Grab key metadata including the creation timestamp, version iteration information, mirror source, and last usage time of the mirror to formulate a scanning strategy;
[0024] S142: Determine the format of the mirror to be scanned by reading the specified identifier and structure information in the mirror file header. When the format of the mirror to be scanned is not the QCOW2 format, and when the mirror format is VMDK or RAW, start the corresponding lossless decompression and format conversion program to convert it to the unified QCOW2 format;
[0025] S143: Store the mirror in the newly created distributed mirror scanning staging file storage, and store the QCOW2 format and the QCOW2 mirror of the processed non-QCOW2 format in the newly created distributed file storage location;
[0026] S144: When the operations of mirror acquisition, decompression, and conversion are completed and stored in the standard mirror cache area, start the integrity check of the system. For QCOW2 format mirrors, perform hash calculation based on the mirrors in the snapshot storage. For non-QCOW2 format mirrors, perform hash calculation based on the QCOW2 format mirrors converted in the mirror cache area, and store the calculated hash values as reference values in the database;
[0027] S145: After performing an operation on the image, calculate the hash value of the image, compare and verify the hash value of the current image with the pre-stored standard hash value. If the two are consistent, it is determined that the image data is true and complete. If they are inconsistent, trigger the alarm mechanism and send an alarm notification containing the detailed information of the image identifier and the error hash value to the operation and maintenance monitoring system.
[0028] Preferably, in step S2, the specific process of the resource scheduling and monitoring module for real-time collecting all-round resource data of the hyper-converged platform and using a preset algorithm to allocate corresponding resource quotas for the scanning tasks is as follows:
[0029] S21: Deploy the Node Exporter and Prometheus servers. Use the Node Exporter to collect various basic resource metrics of the hyper-converged platform nodes, and then aggregate, store, and query the data through the Prometheus server to provide data support for subsequent processes;
[0030] S22: The Prometheus server uses a specified query language to write query statements, summarize the resource data stored in the database at a specified time node, establish a resource abundance scoring mechanism, determine the optimal execution cluster nodes of the scanning tasks according to the resource abundance scoring results, and allocate corresponding CPU computing power for each scanning task according to the node CPU resource status and the complexity of the scanning task, and set the memory capacity and network bandwidth;
[0031] S23: Assign priorities to the tasks according to the task source and task timeliness, create a task queue for all submitted scanning tasks based on the task priorities, and allocate the optimal execution nodes for each task in the task queue. According to the task queue, unload the tasks to the allocated optimal execution nodes for processing in sequence;
[0032] S24: Set multi-dimensional resource monitoring thresholds, monitor the resource occupation during task execution in real time, and at the same time use a preset scheduling algorithm to adjust and update the task scheduling strategy in real time.
[0033] Preferably, the specific process of step S3 is as follows:
[0034] S31: Establish a data connection with the NVD and CVE databases, obtain the latest vulnerability information through a regular data synchronization interface, clean the data obtained from different sources, remove duplicate and invalid information, standardize the unified format, and extract key information including vulnerability numbers, vulnerability descriptions, affected version ranges, and repair suggestions for integrated storage in the local knowledge base to build a unified data structure;
[0035] S32: Set up a real-time monitoring component. When emergency security updates are released in the NVD and CVE databases or major vulnerability disclosures occur in the industry, immediately trigger the update process and update relevant information to the knowledge base within a specified time;
[0036] S33: Through the initial screening by file fingerprint and feature matching, intelligent version comparison to eliminate misjudgments, and a progressive scanning method of key complex component code analysis, comprehensively check for vulnerabilities from coarse-grained to fine-grained;
[0037] S34: Build a malware feature recognition model. Collect malware samples from multiple channels such as public malware sample libraries and sample data shared by security agencies, perform multi-dimensional feature extraction on the samples, construct feature vectors, and then input them into the malware feature recognition model for training;
[0038] S35: For various files in the virtual machine image to be scanned, calculate the file hash value, parse the file header structure, and extract the semantic features of the code segment in turn according to the feature extraction method in the training stage, construct the feature vector of the file to be detected, and input it into the trained malware feature recognition model;
[0039] S36: The malware feature recognition model outputs the malware probability value of each file and compares it with the set probability threshold. When the probability value is greater than the threshold, mark the file as a suspected malware;
[0040] S37: For suspected malware, use intelligent feature matching technology based on a machine learning model for verification. This model is trained based on the typical feature patterns of known malware. By comparing the file features with the typical feature patterns including specified code segment sequences and file structure layouts for matching, if the match is successful, determine that the file is malware, record its location and type information in the image in detail, and generate a malware recognition report.
[0041] Preferably, the specific process of step S4 is as follows:
[0042] S41: Establish an efficient data transmission interface between the scanning engine module and the result processing module. Adopt a communication mechanism based on message queues to enable real-time and stable data transmission, define a standardized data format, and the scanning engine outputs result data according to a unified structure;
[0043] S42: Perform format verification on the received data to determine whether it conforms to the predefined standard format. If not, feedback error information to the scanning engine module, request resending or correction;
[0044] S43: Clean the received data, remove duplicate and redundant data records. For data with some keyword fields missing, fill them reasonably according to the data distribution characteristics and business knowledge. For the missing risk level fields of certain vulnerabilities, infer and fill them in combination with the known risk levels of similar vulnerabilities and the relevant descriptions of the vulnerabilities. For those that cannot be filled, mark them as data to be confirmed and store them separately for subsequent manual processing;
[0045] S44: Extract valuable features from the original data. For vulnerability data, extract vulnerability types, affected software components, and characteristics of the scanning tools that discovered the vulnerabilities; for malware data, extract file hash values, propagation path characteristics, infected system areas, etc.; for configuration defect data, extract configuration file types, involved user privilege levels, security policy category characteristics, and combine the extracted features into a structured feature vector for subsequent algorithms to use;
[0046] S45: Divide the extracted feature vectors into major categories of vulnerabilities, malware, and configuration defects, and further subdivide them under each major category. For example, divide vulnerabilities into subcategories of operating system vulnerabilities, application software vulnerabilities, and network protocol vulnerabilities; divide malware into subcategories of worms, Trojans, and viruses; divide configuration defects into subcategories of improper user privilege configuration, encryption policy defects, and audit policy defects, forming a well - defined classification system;
[0047] S46: Find combinations of vulnerabilities and malware that occur simultaneously, and analyze whether there is a situation where malware uses specific vulnerabilities to spread or attack; explore the associations between configuration defects and vulnerabilities, malware, including whether overly high user privilege configurations lead to being more vulnerable to specific types of malware attacks or vulnerability exploitations; generate an association analysis report through association analysis.
[0048] Preferably, after step S46, the following process is further included:
[0049] S47: Calculate basic statistics such as the number and occurrence frequency of various vulnerabilities, malware, and configuration defects, and create a risk rating model;
[0050] S48: For each problem discovered by scanning, extract the relevant feature information of its corresponding vulnerabilities, malware, and configuration defects, substitute them into the risk rating model for calculation to obtain the risk rating result, associate and store the risk rating result with the problem details, and clearly mark the risk level of each problem in the database;
[0051] S49: Configure multiple report templates in the template engine, define dynamic data filling areas in the template engine, associate them with the scan result data and risk rating information in the database, and based on the open-source visualization chart library, convert various previously generated visualization charts into the form of pictures or HTML fragments, embed them in the corresponding positions of the scan report, and generate targeted repair suggestions for vulnerabilities, malware, and configuration defects.
[0052] In a second aspect, a static virtual machine image scanning system based on a hyper-converged platform is provided for implementing the static virtual machine image scanning method based on a hyper-converged platform, including an image acquisition and preprocessing module, a resource scheduling and monitoring module, a comprehensive scanning engine module, and a result processing and report generation module. The image acquisition and preprocessing module is connected to the resource scheduling and monitoring module, the resource scheduling and monitoring module is connected to the comprehensive scanning engine module, and the comprehensive scanning engine module is connected to the result processing and report generation module;
[0053] The image acquisition and preprocessing module is used to create a consistent snapshot of the storage of the image to be scanned, obtain the static virtual machine image file based on the hyper-converged distributed nodes, use the indexing and retrieval algorithm, extract and record its metadata information, perform multiple rounds of hash calculation on the virtual machine image file, and compare and verify it with the preset standard hash value;
[0054] The resource scheduling and monitoring module is used to collect all-round resource data of the hyper-converged platform in real time, and use the preset algorithm to allocate corresponding resource quotas for the scan task, including the optimal execution node of the scan task, the allocated number of CPU cores, the set memory capacity, and the network bandwidth limit;
[0055] The comprehensive scanning engine module is used to perform multi-stage in-depth scanning. First, it performs a preliminary screening by file fingerprint and feature matching, then excludes false positives through intelligent version comparison, and finally uses static code analysis to deeply dig out potential high-risk vulnerabilities for key complex components. It quickly identifies and locates known malware by extracting and analyzing multi-dimensional features of various files in the virtual machine image, and comprehensively reviews the configuration files in the image;
[0056] The result processing and report generation module is used to comprehensively collect various result data generated by the scanning engine module during the scanning process, and use the preset algorithm to deeply classify, organize, correlate, analyze, and summarize and statistically analyze various result data.
[0057] The beneficial effects of the present invention include:
[0058] The virtual machine image static scanning method and system based on a hyper-converged platform provided by the present invention create a storage snapshot of the image to be scanned, use indexing and retrieval algorithms to obtain static virtual machine image files, extract and record their metadata information, and then perform integrity verification; collect all-round resource data of the hyper-converged platform in real time, and use a preset algorithm to allocate corresponding resource quotas for scanning tasks; perform multi-stage in-depth scanning. First, perform preliminary screening by file fingerprint and feature matching, then exclude misjudgments through intelligent version comparison, and use static code analysis for key complex components to deeply dig potential high-risk vulnerabilities. Quickly identify and locate known malware, and conduct a comprehensive review of the configuration files in the image; use a preset algorithm to deeply classify, correlate, and summarize various result data, achieving virtual machine image scanning in a large-scale hyper-converged cluster in a short time.
[0059] First, the architecture design of the virtual machine image static scanning for the deeply integrated hyper-converged platform covers the technical implementation of seamless docking and collaborative work in each link of image acquisition, comprehensive scanning, resource scheduling, and result feedback with the hyper-converged platform, giving full play to the distributed, elastic, and integrated characteristics of the hyper-converged platform. By optimizing the image parsing process and intelligently scheduling resources, the present invention can complete virtual machine image scanning in a large-scale hyper-converged cluster in a short time, greatly shortening the waiting time.
[0060] Second, through the construction of a multi-source fusion vulnerability knowledge base and multi-stage in-depth scanning technology, static malware detection, and fine-grained configuration compliance detection in the comprehensive scanning engine module, it can quickly and efficiently identify malware, etc., effectively reducing risks.
[0061] Finally, through the intelligent prediction based on real-time resource data of the hyper-converged platform, the optimization algorithm of the dynamic allocation strategy, and the abnormal adjustment mechanism in the resource scheduling and optimization module, fast emergency handling is achieved. The dynamic resource scheduling strategy ensures that the scanning task will not cause too much burden on the normal operation of the hyper-converged platform, maintaining a good performance level. With the help of an advanced false alarm filtering mechanism, the present invention effectively reduces the number of false alarms, enabling administrators to focus more on truly important security issues. Brief Description of the Drawings
[0062] Figure 1 It is a structural diagram of the virtual machine image static scanning system based on the hyper-converged platform of the present invention.
[0063] Figure 2 It is a schematic flowchart of the virtual machine image static scanning method based on the hyper-converged platform of the present invention.
[0064] Figure 3 It is a flowchart of image acquisition and image preprocessing of the present invention.
[0065] Figure 4It is the flowchart of the resource scheduling and monitoring module of the present invention.
[0066] Figure 5 It is the flowchart of the comprehensive scanning engine module of the present invention. Detailed implementation manners
[0067] The following will Figures 1 to 5 make a further detailed description of the present invention:
[0068] Embodiment 1
[0069] Refer to the Figures 1 - 5 shown. A static scanning method for virtual machine images based on a hyper-converged platform includes the following steps:
[0070] S1: Create a consistent snapshot of the storage of the image to be scanned through the image acquisition and preprocessing module. Based on the hyper-converged distributed nodes, use the indexing and retrieval algorithm to obtain the static virtual machine image file, extract and record its metadata information, and perform multiple rounds of hash calculation on the virtual machine image file, and compare and verify it with the preset standard hash value. Deeply integrate the distributed storage architecture of the hyper-converged platform. When starting the image acquisition process, first create a consistent snapshot of the storage of the image to be scanned to ensure the consistency and integrity of the data, and avoid affecting the scanning result due to subsequent operations or dynamic changes during system operation. After the snapshot is generated, based on the hyper-converged distributed nodes, use the indexing and retrieval algorithm to achieve the agile acquisition of a large number of static virtual machine images. While obtaining the image, extract and record its metadata information, such as the creation timestamp of the image, version iteration information, etc., to provide multi-dimensional decision-making basis for the subsequent scanning process.
[0071] S2: Through the resource scheduling and monitoring module, collect all-round resource data of the hyper-converged platform in real time, and use the preset algorithm to allocate corresponding resource quotas for the scanning task, including the optimal execution node of the scanning task, the allocated number of CPU cores, the set memory capacity, and the network bandwidth limit. The resource scheduling and monitoring module fully considers the distributed characteristics of the hyper-converged platform and the needs of executing parallel scanning tasks, collects all-round resource data of the hyper-converged platform in real time, and conducts all-round and dynamic monitoring and evaluation of the real-time utilization rate of the computing resources (CPU, memory) of the platform nodes, the storage I / O load status, and the network bandwidth utilization rate. Based on the above data, allocate the most suitable resource quotas for the scanning task, including determining the optimal execution node of the scanning task, allocating appropriate CPU cores, setting the memory capacity, and the network bandwidth limit and other key parameters. At the same time, the module can intelligently arrange multiple image scanning tasks to be executed in parallel according to the real-time usage of the platform resources. Through comprehensive consideration of various factors such as the resource requirements of the task and the estimated execution time, formulate the optimal parallel execution plan to ensure that the resources of the hyper-converged platform are utilized to the maximum extent without affecting the system stability, and accelerate the scanning process.
[0072] S3: Conduct multi-stage in-depth scanning through the comprehensive scanning engine module. First, conduct a preliminary screening using file fingerprints and feature matching. Then, exclude false positives through intelligent version comparison. Finally, conduct in-depth exploration of potential high-risk vulnerabilities for critical and complex components using static code analysis. Conduct multi-dimensional feature extraction and analysis of various files in the virtual machine image for rapid identification and location of known malware, and conduct a comprehensive review of the configuration files in the image;
[0073] S4: Through the result processing and report generation module, comprehensively collect various result data generated by the scanning engine module during the scanning process, and use preset algorithms to conduct in-depth classification, correlation analysis, and summary statistics on various result data.
[0074] In this embodiment, comprehensively collect various result data generated by the scanning engine module during the scanning process, and use data mining and machine learning algorithms to conduct in-depth classification, correlation analysis, and summary statistics on the result data. Based on multi-dimensional factors such as the type and harm degree of vulnerabilities, the propagation characteristics and destruction potential of malware, and the influence scope and repair difficulty of configuration defects, construct a risk rating model, and use this model to conduct accurate risk assessment and level classification for each problem found in the scan. For example, rate high-risk vulnerabilities that can directly lead to remote code execution, complete loss of system permissions, or large-scale leakage of sensitive data as the highest risk level, and rate configuration problems that may only cause minor information leakage or have a minor impact on system performance as a low risk level.
[0075] Based on the scan result data and risk rating information, use the template engine and data visualization technology to generate a detailed and intuitive scan report. The report content not only covers the basic information of the virtual machine image (such as name, unique identifier, etc.), the start time and end time of the scan task, the details of the scan results (detailed classification and listing according to categories such as vulnerabilities, malware, configuration defects, etc.), the risk rating results, and professional repair suggestions for each specific problem, but also provides value-added information such as experience reference based on historical repair cases of similar problems and predictive analysis of the effect after repair. The scan report supports generation and output in multiple common formats (such as HTML, PDF, etc.) to meet the diverse needs of different user groups and application scenarios. At the same time, this module also has a function of exporting result data, which can seamlessly export the scan result data to other security management systems, work order management systems, or big data analysis platforms within the enterprise for subsequent repair process tracking and monitoring, security situation analysis, and data mining and utilization operations.
[0076] Embodiment 2
[0077] Based on Embodiment 1, the specific process of creating a consistent snapshot for storing the mirror to be scanned through the mirror acquisition and preprocessing module in step S1 is as follows:
[0078] S11: The mirror acquisition and preprocessing module receives an external start instruction or is activated at a specified time node to start executing tasks;
[0079] S12: By calling the snapshot creation interface provided by the distributed storage system, a consistent snapshot is generated for the target mirror storage when the acquisition is started;
[0080] S13: Based on the hyper-converged distributed nodes, set up mirror indexing and retrieval, quickly locate and obtain the target mirror in the massive static virtual machine images, obtain the mirror metadata information from the distributed file mirror storage snapshot storage, screen out the specified index items from the mirror metadata information obtained from the snapshot storage, and use a preset hash algorithm to allocate the node ID and the storage location of the index item;
[0081] S14: Create a distributed mirror scanning staging file to store the mirror to be scanned, obtain the list of mirrors to be scanned, record the metadata information of the mirrors to be scanned, judge the format of the mirrors to be scanned, store the mirrors in the newly created distributed mirror scanning staging file, and perform mirror integrity calculation.
[0082] In this embodiment, the specific process of step S13 is as follows:
[0083] S131: Obtain the mirror metadata information from the distributed file mirror storage snapshot storage, and screen out the key features including the mirror name, creation time, and size as index items;
[0084] S132: Use a preset hash algorithm to allocate the node ID and the storage location of the index item. Each node generates a hash value as the node ID based on the comprehensive information of its own network address and hardware identifier;
[0085] S133: For the index item, calculate its hash value, and map the index item to the corresponding node through the preset hash algorithm;
[0086] S134: When a retrieval request is received, according to the given query conditions, multi-node parallel retrieval is performed to quickly locate the target mirror index.
[0087] The specific process of step S14 is as follows:
[0088] S141: Create a distributed mirror scanning staging file to store the mirror to be scanned, obtain the list of mirrors to be scanned, capture the key metadata including the creation timestamp, version iteration information, mirror source, and last use time of the mirror to formulate a scanning strategy;
[0089] S142: Determine the format of the image to be scanned by reading the specified identification and structure information in the image file header. When the format of the image to be scanned is not the QCOW2 format, and when the image format is VMDK or RAW, start the corresponding lossless decompression and format conversion program to convert it to the unified QCOW2 format;
[0090] S143: Store the image in the newly created distributed image scanning temporary file storage, and store the QCOW2 format and the QCOW2 image of the processed non-QCOW2 format in the newly created distributed file storage location;
[0091] S144: When the operations of image acquisition, decompression, and conversion are completed and stored in the standard image buffer, start the integrity check of the system. For the QCOW2 format image, calculate the hash value based on the image in the snapshot storage. For the non-QCOW2 format image, calculate the hash value based on the QCOW2 format image after conversion in the image buffer, and store the calculated hash value as the reference value in the database;
[0092] S145: After performing a round of operations on the image, calculate the hash value of the image, and compare and verify the hash value of the current image with the pre-stored standard hash value. If the two are consistent, it is determined that the image data is true and complete. If they are inconsistent, trigger the alarm mechanism and send an alarm notification containing the detailed information of the image identifier and the error hash value to the operation and maintenance monitoring system.
[0093] In this embodiment, first initiate the start instruction of the entire module operation to trigger a series of subsequent acquisition and preprocessing operations for the virtual machine image. The system receives external instructions or activates the image acquisition and preprocessing module to start executing tasks according to the preset timing tasks.
[0094] With the help of the distributed file of the hyper-converged platform, generate a consistent snapshot for the target image storage during startup acquisition. Use copy-on-write to ensure that at the moment of creating the snapshot, the image data is completely "frozen". Subsequent operations such as changes in the processing flow of the module itself or dynamic updates and modifications during the system operation will not interfere with the consistency and integrity of the original image data, thus providing a reliable data basis for subsequent scans and ensuring the authenticity of the scan results. Trigger the snapshot generation operation by calling the snapshot creation interface provided by the distributed storage system. Set up image indexing and retrieval based on the hyper-converged distributed nodes. Utilize the constructed indexing and retrieval system to quickly locate and obtain the target image in the massive static virtual machine images, shorten the acquisition time, improve the overall efficiency, and ensure the timely progress of the subsequent processes.
[0095] The image acquisition and preprocessing module mainly includes a distributed index construction module, an image retrieval module, a consistency algorithm module, and a Bloom filter description.
[0096] The distributed index module is constructed to obtain mirror metadata information from the distributed file mirror storage snapshot storage for easy retrieval. Key features such as mirror name, creation time, size, etc. are selected as index items. The consistent hashing algorithm is used to allocate node IDs and the storage locations of index items. Each node generates a hash value as the node ID based on comprehensive information such as its network address and hardware identifier. For index items, their hash values are also calculated, and the index items are mapped to the corresponding nodes through the consistent hashing algorithm. Compared with the traditional modulo operation, this method can effectively reduce the amount of data migration and ensure the stability of the system when nodes are dynamically added or removed. For example, when a new node is added, only a small number of index items need to be redistributed.
[0097] The mirror retrieval module performs parallel retrieval on multiple nodes according to the given query conditions to quickly locate the target mirror index. When a retrieval request is received, whether it is an exact search by mirror name or a filter by time period, etc., the system first searches in the global index cache. The global index cache stores recently frequently accessed index items and the corresponding node location information. If a hit occurs, a retrieval request is directly sent to the corresponding node, greatly accelerating the initial response speed. After receiving the retrieval request, each node uses the local Bloom filter to quickly exclude the index areas where the target mirror definitely does not exist. The Bloom filter is constructed based on the hash values of index items and can efficiently judge whether an element is in the set with extremely small memory occupancy, avoiding invalid retrievals. Then, the node compares and filters according to the index items stored locally against the retrieval conditions to find the set of mirrors that meet the conditions, and finally aggregates the results of each node to obtain the retrieval result set.
[0098] The consistency algorithm module, taking the mirror name as an example, constructs a consistent hashing ring of 0 to 2^32 - 1: Assume that the hyper-converged platform contains 4 nodes, and the node identifiers are HNode1, HNode2, HNode3, and HNode4 respectively, and their corresponding network addresses are "192.168.12.101", "192.168.12.102", "192.168.12.103", and "192.168.12.104".
[0099] Use the MurmurHash function to calculate the hash of these node identifiers and map the results to the range of 0 to 2^32 - 1. For HNode1, MurmurHash("192.168.12.101") after conversion (converting the binary sequence output by the MurmurHash function to the range of 0 to 2^32 - 1 according to the unsigned integer rule, and the accuracy is ensured by the underlying logic of the function), assume the hash value obtained is 35000000. Similarly, for HNode2, "192.168.12.102" after calculation is assumed to get a hash value of 80000000, for HNode3 it is 125000000, and for HNode4 it is 180000000. Arrange these node hash values in ascending order in a virtual circular space, and the range of this circular space is 0 to 2^32 - 1. In this way, a basic consistent hash ring is constructed, and each server node occupies a specific position on the ring, providing a basis for subsequent data storage allocation.
[0100] Allocate index entries to nodes: There are a large number of virtual machine images in the system that need to be stored and managed, and these images use names as the key identifiers. For example, there are several typical image names: "webserver_vm_001.qcow2", "dbserver_vm_002.vmdk", "appserver_vm_003.raw", etc. Also use the MurmurHash function to calculate the hash of these image names and map them to the range of 0 to 2^32 - 1. Assume that after calculation and conversion by the MurmurHash function, "webserver_vm_001.qcow2" gets a hash value of 60000000, "dbserver_vm_002.vmdk" gets a hash value of 100000000, and "appserver_vm_003.raw" gets a hash value of 150000000.
[0101] Mapping of Image Name to Hyper-Converged Nodes: According to the rules of the consistent hashing algorithm, starting from the position corresponding to the hash value of each image name on the ring, search clockwise to find the hash value of the first node that is greater than or equal to this hash value. The corresponding hyper-converged platform node is the storage location of the image metadata. For example, for "webserver_vm_001.qcow2", the first node hash value greater than or equal to its hash value 60000000 found clockwise is 80000000 (HNode2). Therefore, the metadata of the "webserver_vm_001.qcow2" image will be stored on HNode2. The hash value of "dbserver_vm_002.vmdk" is 100000000, and the first node hash value greater than or equal to it found clockwise is 125000000 (HNode3). So, its related image metadata is stored on HNode3. The hash value of "appserver_vm_003.raw" is 150000000, and the first node hash value greater than or equal to it found clockwise is 180000000 (HNode4). So, its related image metadata is stored on HNode4.
[0102] Retrieval and Data Acquisition Process Based on Consistent Hashing: Client Retrieval Request Processing: When the client initiates a request to retrieve a specific image, whether the request is for an exact search based on the image name or for filtering based on other associated conditions (as long as these conditions can be associated with a specific image name), the client first needs to perform a hash calculation on the retrieval key information (such as the image name) using the MurmurHash function, and map the result to the range of 0 to 2^32 - 1 as well. For example, if you want to retrieve the "webserver_vm_001.qcow2" image, first calculate its hash value 60000000 (the same as the calculation during storage).
[0103] Locate the Index Storage Node and Obtain Data: According to the established rules of the consistent hashing ring, starting from this hash value, quickly locate the hyper-converged platform node that may store the target image metadata in the clockwise direction. In this example, starting from the position corresponding to 60000000 and searching clockwise, locate HNode2 as the node that may store the target image index. Subsequently, the client sends a data acquisition request to HNode2. HNode2 searches for the metadata of the "webserver_vm_001.qcow2" image in its local storage area and returns the result to the client. Make full use of the efficient characteristics of the consistent hashing algorithm to ensure the rapidity and accuracy of image index retrieval and acquisition, and improve the operating efficiency of the entire distributed system.
[0104] Illustration of Bloom Filter:
[0105] Suppose there is a hyper-converged node that locally stores 1,000 index entries covering various mirror-related information. To construct a Bloom filter, first determine the appropriate number of hash functions (assuming MurmurHash, FNV1a hash, and BKDR hash are selected) and the number of bits of the Bloom filter (assuming 10,000 bits, determined comprehensively based on factors such as data volume and false positive rate).
[0106] For each index entry, such as "webserver_vm_001.qcow2", calculate the hash values using these 3 hash functions respectively:
[0107] MurmurHash("webserver_vm_001.qcow2") = 12345678 (corresponding to the 12,345,678th bit of the Bloom filter, set this bit to 1);
[0108] FNV1a hash("webserver_vm_001.qcow2") = 98765432 (similarly processed, 98765432 % 10000 = 5432, set the 5432nd bit to 1);
[0109] BKDR hash("image_file_001") = 43219876 (43219876 % 10000 = 9876, set the 9876th bit to 1);
[0110] When a retrieval request to search for the mirror named "search_image.qcow2" is received, calculate the hash values of "search_image.qcow2" using these 3 hash functions in the same way:
[0111] MurmurHash("search_image.qcow2") = 65432109;
[0112] FNV1a hash("search_image.qcow2") = 87654321;
[0113] BKDR hash("search_image.qcow2") = 32109876;
[0114] Check the three digits corresponding to 65432109 % 10000 = 2109, 87654321 % 10000 = 4321, and 32109876 % 10000 = 9876 in the Bloom filter. If any one of them is 0, it can be determined that "search_image.qcow2" is definitely not in the locally stored index entries, and it can be directly excluded to avoid ineffective retrieval. Only when all three digits are 1, will it enter the subsequent detailed index comparison and screening process, greatly improving the retrieval efficiency and achieving efficient pre-filtering with extremely small memory occupancy (10,000 bits, approximately 1.2KB).
[0115] Create a distributed mirror scan staging file storage to store the mirrors to be scanned, and obtain the list of mirrors to be scanned, whether the request is for precise search based on the mirror name or screening according to other associated conditions.
[0116] Record the metadata information of the mirrors to be scanned, and capture key metadata such as the creation timestamp of the mirror, version iteration information, the source of the mirror (which source virtual machine it was cloned from), and the last usage time. These information can provide a reference basis for the subsequent scanning process. For example, predict the possible range of vulnerabilities based on version iteration, assist in analyzing the historical evolution of the mirror according to the creation timestamp, trace the source of the problem in combination with the source information, help make the scanning strategy more targeted, and improve the scanning efficiency and accuracy.
[0117] By parsing the metadata storage area inside the mirror file (different formats of mirrors have different metadata storage methods, such as QCOW2 mirrors have relevant records in the file header or specific metadata blocks, and VMDK mirrors also have their fixed metadata storage structure), use the corresponding parsing tools or library functions to read and extract the required metadata information to provide support for subsequent operations, and store it in the local database for subsequent use.
[0118] Judge the format of the mirror to be scanned: Since there are multiple storage formats for virtual machine mirrors, to unify the processing process, it is necessary to accurately identify the mirror format. By reading the specified identification, structure information, etc. in the mirror file header, determine whether it belongs to QCOW2, VMDK, RAW or other formats, and provide a decision basis for subsequent format conversion.
[0119] Start the corresponding lossless decompression and format conversion programs under specified circumstances. When it is judged that the mirror is not in QCOW2 format, such as in VMDK format, start the corresponding lossless decompression and format conversion programs (with the help of tools such as qemu-img), and convert it to a unified QCOW2 format to ensure that the subsequent scanning process does not need to be separately adapted for different formats, simplify the process, improve compatibility, and ensure the smooth progress of the scanning work.
[0120] Taking the conversion using the qemu-img tool as an example, execute the following command under the Linux system to complete the conversion: qemu-img convert -f vmdk -O qcow2 input.vmdk output.qcow2;
[0121] Store the image in the newly created distributed image scanning staging file storage, and place the qcow2 format and processed non-qcow2 images in the newly created distributed file storage location for subsequent scanning processes to achieve centralized management and efficient utilization of image data.
[0122] Image integrity calculation: To prevent image data corruption caused by accidental errors during storage and transmission or scanning processes, when operations such as image acquisition, decompression, and conversion are completed and stored in the standard image buffer, the integrity verification module of the system is started. For qcow2 format images, hash calculation is performed based on the images in the snapshot storage. For non-qcow2 format images, hash algorithm calculation is performed based on the qcow2 format images converted in the image buffer. At the same time, the calculated hash value is stored in the database as the reference value.
[0123] After a round of operations (such as scanning) on the image, the hash value of the image needs to be calculated. Finally, the hash value of the current image is compared with the pre-stored standard hash value for verification. If the two match exactly, it indicates that the image data is true and complete. If there are differences, the alarm mechanism is triggered, and an alarm notification containing detailed information such as the image identifier and the incorrect hash value is sent to the operation and maintenance monitoring system.
[0124] Implementation logic of the hash algorithm:
[0125] In the first round, calculate the hash value for each 4KB data block one by one. The formula is: H = sha256(Mi), where H is the hash value of the data block Mi, and i represents the i-th data block. Use the hash calculation library function (such as the sha256 calculation function provided by OpenSSL, and pass the memory address of the data block as a parameter when calling this function) to pass the memory address of each data block into the function to obtain the corresponding hash value. Subsequently, concatenate the hash values of all data blocks into a string in sequence, and execute the SHA-256 calculation again to obtain the final image hash value Hfinal.
[0126] Embodiment 3
[0127] Based on Embodiment 1 or Embodiment 2, the specific process of the resource scheduling and monitoring module in step S2 for real-time collecting all-round resource data of the hyper-converged platform and allocating corresponding resource quotas for the scanning task using a preset algorithm is as follows:
[0128] S21: Deploy the Node Exporter and Prometheus servers. Use the Node Exporter to collect various basic resource metrics of the nodes in the hyper-converged platform, and then aggregate, store, and query the data through the Prometheus server to provide data support for subsequent processes;
[0129] S22: The Prometheus server uses a specified query language to write query statements to summarize the resource data stored in the database at a specified time node, establish a resource adequacy scoring mechanism, determine the optimal execution cluster node for the scanning task based on the resource adequacy scoring result, and allocate corresponding CPU computing power for each scanning task according to the node's CPU resource status and the complexity of the scanning task, and set the memory capacity and network bandwidth;
[0130] S23: Assign priorities to tasks according to the task source and task timeliness, create a task queue for all submitted scanning tasks based on the task priorities, and allocate the optimal execution node for each task in the task queue. According to the task queue, unload the tasks to the allocated optimal execution nodes for processing in sequence;
[0131] S24: Set multi-dimensional resource monitoring thresholds, monitor the resource occupancy during task execution in real time, and at the same time use a preset scheduling algorithm to adjust and update the task scheduling strategy in real time.
[0132] Deploy the Node Exporter and Prometheus servers: Use the Node Exporter to collect various basic resource metrics of the nodes in the hyper-converged platform, and then aggregate, store, and query the data through Prometheus to provide data support for subsequent processes. Install NodeExporter on each node of the hyper-converged platform and set it to start automatically at boot so that it can continuously collect node resource data. Install and configure the Prometheus server, and specify the target nodes to be monitored in the configuration file, that is, all cluster nodes of the hyper-converged platform. Start the Prometheus server, and it will send data collection requests to the Node Exporter of each node at the configured time interval (default 15 seconds) to obtain resource data.
[0133] Collection Node Computing Resource Utilization Rate: Understand the usage of CPU and memory to judge the remaining resources later and allocate scanning task nodes reasonably. Prometheus collects CPU utilization rate metrics by sending requests to Node Exporter. In the query interface of Prometheus or through its API, use the node_cpu_usage_seconds_total metric to obtain the cumulative CPU usage time of each node. Combine information such as the system startup time to calculate the current CPU utilization rate. Store the obtained data in the time series database built into Prometheus in the form of time series. The record format is similar to: {Node ID: [Timestamp 1, CPU Utilization Rate 1], [Timestamp 2, CPU Utilization Rate 2],...}, where the Node ID can correspond to the IP address or custom name of the node for subsequent identification.
[0134] For memory utilization rate, use the node_memory_MemTotal_bytes and node_memory_MemUsed_bytes metrics to obtain the total capacity and used amount of the node's memory respectively. Calculate the real-time utilization rate through simple calculations and store it in the Prometheus database in a similar format: {Node ID: [Timestamp 1, Memory Utilization Rate 1], [Timestamp 2, Memory Utilization Rate 2],...}.
[0135] Monitor Storage I / O Load Conditions: Understand the busyness of storage devices, avoid delays in scanning tasks due to storage I / O bottlenecks, and ensure efficient task execution. Node Exporter cooperates with storage driver plugins to collect storage I / O metrics. Prometheus obtains metric data such as node_disk_io_time_seconds_total (cumulative disk I / O operation time), node_disk_read_bytes_total (cumulative disk read bytes), and node_disk_write_bytes_total (cumulative disk write bytes) through the configured collection rules. Classify and store them in the Prometheus database according to nodes and storage devices. For example: {Node ID: {Storage Device ID 1: [Timestamp 1, IOPS 1, Read / Write Bandwidth 1], [Timestamp 2, IOPS 2, Read / Write Bandwidth 2],...}, Storage Device ID 2: [...]}}, where metrics such as IOPS can be obtained through further calculations of relevant metric data, such as estimating the read / write bandwidth based on the number of read / write bytes and time interval.
[0136] Tracking network bandwidth utilization: Ensure that the network resources allocated for scanning tasks are reasonable, without affecting the network communication of other platform services, and at the same time meeting the data transmission requirements of scanning tasks. Node Exporter uses network interface statistics to collect network inbound and outbound traffic data, and Prometheus obtains relevant metrics such as node_network_receive_bytes_total (cumulative network received bytes) and node_network_transmit_bytes_total (cumulative network transmitted bytes). Combining with the bandwidth configuration information of the node network interface (which can be pre-configured in Prometheus or obtained from other configuration management systems), calculate the network bandwidth utilization. The storage format is: {node ID: [timestamp 1, network bandwidth utilization 1], [timestamp 2, network bandwidth utilization 2],...}, and it is stored in the Prometheus database.
[0137] Evaluation and task resource allocation based on the collected data, data summary and analysis:
[0138] Prometheus provides a powerful query language PromQL. Use it to write query statements to summarize the resource data stored in the database at regular intervals (such as every 1 minute). For example, calculate the average CPU usage rate of each node in the past 5 minutes through statements similar to avg_over_time(node_cpu_usage_seconds_total[5m]). Similarly, key statistics such as the peak value, average value of storage I / O load, and average value of network bandwidth utilization can be calculated. Store these summary data in a database table dedicated for task scheduling for subsequent quick access.
[0139] Determine the optimal execution cluster nodes for scanning tasks:
[0140] Find the nodes with relatively abundant resources and most suitable for performing scanning tasks, improve the task execution efficiency, and reduce resource competition. Set the scoring rules for resource abundance. For example, assign certain weights (such as 0.3, 0.3, 0.2, 0.2) to the CPU, memory, storage I / O, and network bandwidth of each node respectively, and score according to the resource statistics of each node obtained through summary analysis. Use PromQL to calculate the node score in combination with the weights. The calculation formula can be: Node score = CPU average utilization weight × (1 - CPU average utilization) + memory average utilization weight × (1 - memory average utilization) + storage I / O load average weight × (1 / (1 + storage I / O load average)) + network bandwidth utilization average weight × (1 - network bandwidth average utilization). The real-time data such as the average utilization rate and other indicators here are obtained through PromQL queries and substituted into the calculation. Sort the nodes according to the scores, and select several nodes with the highest scores (determined according to the number of scanning tasks) as candidate execution nodes.
[0141] Allocate CPU cores: According to the CPU resource status of the node and the complexity of the scanning task, allocate appropriate CPU computing power for each scanning task to ensure the smooth progress of the task. Estimate the CPU requirements of different types of scanning tasks, and establish a mapping table between the task type and the CPU core number requirements. For example, set 2 cores for simple file scanning tasks and 4 cores for complex system vulnerability scanning tasks, etc.
[0142] Set the memory capacity: Provide sufficient memory space for the scanning task to prevent task lag or failure due to insufficient memory, and at the same time avoid memory waste. Similar to the CPU core allocation, estimate the memory requirements based on the scanning task type. For example, set 4GB of memory for small image scanning and 8GB of memory for large and complex image scanning, etc., to form a memory requirement mapping table. Consider the node memory utilization rate. If the current node memory utilization rate is lower than 50%, allocate the full amount according to the requirements; if it is higher than 80%, then appropriately compress the allocation amount (such as allocate 60% of the requirements), and record the allocation situation: {scanning task ID: [execution node ID, allocated memory capacity]}.
[0143] Set network bandwidth limits: While meeting the data transfer requirements of scanning tasks, ensure the network smoothness of other platform services and avoid network congestion. Analyze the data transfer characteristics of scanning tasks and estimate the required network bandwidth, such as estimating based on factors like image size. For regular image scanning tasks, set a 100Mbps bandwidth; for in-depth scans with large amounts of data, set a 500Mbps bandwidth. Considering the current network bandwidth utilization rate of the node, if the utilization rate is below 40%, allocate the full estimated bandwidth; if it is above 60%, reduce it by a certain proportion (such as 70% of the estimated bandwidth), and record: {Scanning task ID: [Execution node ID, Allocated network bandwidth]}. At the network level, utilize the software-defined network (SDN) function of the hyper-converged platform to limit the network traffic of scanning tasks by setting traffic policies or QoS (Quality of Service) rules to ensure that the bandwidth allocation meets the requirements.
[0144] Intelligent parallel task scheduling, task queuing, and priority setting: All submitted scanning tasks enter the task queue. The system assigns a priority value (such as 1 - 10, with 10 being the highest priority) to each task based on factors such as the task source (e.g., system critical update scanning tasks have a high priority, while daily scanning tasks initiated by ordinary users have a low priority) and task timeliness requirements (urgent tasks have a high priority). The task queue is sorted from highest to lowest priority, and the task submission time is also recorded to ensure that in the case of the same priority, tasks are executed in the order of first come, first served. Use a memory queue structure to store task queue information to ensure efficient enqueueing and dequeueing operations.
[0145] Formulate a parallel execution plan: Starting from the head of the task queue, select tasks in sequence and formulate a detailed execution schedule in combination with the determined optimal execution nodes and resource allocation. For example, Task 1 starts execution on Node A at 10:00 and is expected to take 30 minutes; Task 2 starts execution on Node B at 10:05 and is expected to take 40 minutes, etc., forming an execution plan list: {Scanning task ID: [Execution node ID, Start time, Expected execution time]}. Through cooperation with the task scheduling system, send the execution plan to each execution node to ensure that tasks start on time. When arranging tasks, fully consider the parallel load capacity of node resources to avoid over-concentration of tasks on the same node leading to resource exhaustion. Optimize the execution plan by simulating the execution of different task combinations on nodes to ensure maximum resource utilization. Some scheduling algorithm simulation tools or analysis models based on historical data can be used to assist in decision-making and continuously adjust the task arrangement strategy.
[0146] Resource Monitoring and Dynamic Adjustment, Real-time Resource Consumption Tracking: Continuously monitor the resource occupancy during the execution of the scanning task, promptly detect potential problems, and provide data support for dynamic adjustment. At set time intervals (such as every 10 seconds), query the Prometheus database again using PromQL to obtain the actual occupancy of CPU, memory, storage I / O, and network bandwidth of the scanning tasks on each execution node, and compare it with the allocated resource quotas. Update the real-time resource consumption data to the task execution status monitoring table in the format: {Scanning Task ID: [Execution Node ID, Current CPU Usage Rate, Current Memory Usage Rate, Current Storage I / O Load, Current Network Bandwidth Utilization Rate]}. This monitoring table is a data structure in memory, used to display the task resource consumption status in real-time, facilitating subsequent early warning judgments.
[0147] Threshold Judgment and Early Warning: Detect abnormal resource consumption in advance and take timely measures to prevent impacts on the normal operations of the platform. Set multi-dimensional resource monitoring thresholds, such as a CPU usage rate threshold of 90%, a memory usage rate threshold of 95%, a peak storage I / O load threshold of 1000 read / write operations per second, a network bandwidth utilization rate threshold of 90%, etc. When it is monitored that a certain resource occupancy indicator of a scanning task exceeds the corresponding threshold, immediately trigger the early warning mechanism and send an alarm message to the system administrator, including detailed information such as the task ID, the type of over-standard resource, and the current resource occupancy value. By integrating with a monitoring system (such as Prometheus's own Alertmanager) and configuring alarm rules, automated alarm notifications can be achieved to ensure that administrators can respond in a timely manner.
[0148] Resource Dynamic Adjustment Strategy: If the CPU usage exceeds the standard, for non-critical tasks that can be paused, suspend their execution to release CPU resources. For running critical tasks, try to transfer part of their computing load to other idle nodes to achieve dynamic task migration through a distributed computing framework; when memory is insufficient, first try to reclaim the cache resources occupied by tasks. If it is still insufficient, suspend some low-priority tasks in the order of priority to release memory. At the same time, for problems such as memory leaks, start a memory optimization program for repair. By combining with the memory management tool jemalloc, detect the memory usage situation. For tasks suspected of memory leaks, conduct memory analysis and optimization, or directly restart the relevant tasks to release memory; when the storage I / O load is too high, adjust the storage access strategy of the scanning task, such as increasing the data read cache, optimizing the data write order, and reducing frequent small data block reads and writes. If the problem is serious, suspend the relevant tasks and wait for the storage load to ease; when the network bandwidth is tight, reduce the data transmission frequency of the scanning task, such as extending the data synchronization interval, giving priority to ensuring the network requirements of critical services, and suspending the network transmission of non-urgent tasks if necessary. By modifying the configuration parameters of the scanning task (such as the transmission interval time) or adjusting the QoS rules at the network level, limit the network bandwidth occupancy of non-critical tasks to ensure the smoothness of the critical service network.
[0149] Fault Troubleshooting and Optimization Repair: Analyze the root cause of resource anomalies, solve problems, and improve system stability and task execution efficiency. When a resource anomaly triggers an alarm, the system automatically starts a fault troubleshooting program to collect relevant log information (including task execution logs, system resource monitoring logs, node hardware logs, etc.), and use data analysis tools to locate the root cause of the problem, such as whether there is a program deadlock resulting in high CPU occupancy, memory leak, storage device failure causing I / O anomalies, network congestion nodes, etc. Some log analysis platforms (such as ELK Stack) can be used to quickly retrieve and analyze the logs to find problem clues. According to the troubleshooting results, take targeted optimization and repair measures. For software problems, such as updating defective code modules and adjusting task configuration parameters; for hardware problems, promptly notify the operation and maintenance personnel to replace or repair the hardware, and adjust the task allocation strategy to avoid faulty nodes. Throughout the process, a detailed problem tracking record should be established for subsequent review and continuous optimization of system performance.
[0150] Embodiment 4
[0151] Based on Embodiment 1 or Embodiment 2 or Embodiment 3, the specific process of step S3 is as follows:
[0152] S31: Establish a data connection with the NVD and CVE databases, obtain the latest vulnerability information through the regular data synchronization interface, clean the data obtained from different sources, remove duplicate and invalid information, standardize the unified format, and extract key information including vulnerability numbers, vulnerability descriptions, affected version ranges, and repair suggestions for integrated storage in the local knowledge base to build a unified data structure;
[0153] S32: Set up a real-time monitoring component. When the NVD and CVE databases release emergency security updates or major vulnerability disclosures occur in the industry, immediately trigger the update process and update the relevant information to the knowledge base within the specified time;
[0154] S33: Conduct a comprehensive vulnerability check from coarse-grained to fine-grained through a progressive scanning method of initial screening by file fingerprint and feature matching, intelligent version comparison to exclude misjudgments, and key complex component code analysis;
[0155] S34: Build a malware feature recognition model. Collect malware samples from multiple channels such as publicly available malware sample libraries and sample data shared by security agencies, extract multi-dimensional features from the samples, and input the constructed feature vectors into the malware feature recognition model for training;
[0156] S35: For various files in the virtual machine image to be scanned, calculate the file hash value, parse the file header structure, and extract the semantic features of the code segment in sequence according to the feature extraction method in the training stage, construct the feature vector of the file to be detected, and input it into the trained malware feature recognition model;
[0157] S36: The malware feature recognition model outputs the malware probability value of each file and compares it with the set probability threshold. When the probability value is greater than the threshold, mark the file as a suspected malware;
[0158] S37: For suspected malware, use the intelligent feature matching technology based on the machine learning model for verification. This model is trained based on the typical feature patterns of known malware. Match the file features with the typical feature patterns including the specified code segment sequence and file structure layout. If the match is successful, determine that the file is malware, record its location and type information in the image in detail, and generate a malware recognition report.
[0159] In this embodiment, the comprehensive scanning engine module includes a vulnerability detection sub-module, a malware identification sub-module, and a configuration compliance review sub-module. The vulnerability detection sub-module creates a multi-source integrated and real-time updated vulnerability knowledge base, including authoritative databases such as NVD and CVE, as well as cutting-edge industry research. Through continuous data collection and analysis, it ensures that the knowledge base always keeps up with the latest trends in the security field, providing solid data support for vulnerability scanning. At the same time, multi-stage in-depth scanning is implemented: first, initial screening is carried out by file fingerprints and feature matching, then version intelligent comparison is used to eliminate misjudgments, and finally, static code analysis (such as symbolic execution and taint analysis) is used for key complex components to deeply dig potential high-risk vulnerabilities. The malware identification sub-module: uses deep learning algorithms to conduct in-depth training and learning on a large number of malware samples, and constructs a highly accurate malware feature identification model with strong generalization ability. Through multi-dimensional feature extraction and analysis of various files in the virtual machine image, including but not limited to file hash value calculation, file header structure parsing, code segment semantic feature extraction, and intelligent feature matching based on machine learning models, etc., rapid identification and positioning of known malware are achieved.
[0160] The configuration compliance review sub-module: deeply analyzes various configuration files in the virtual machine image, and adopts syntax-semantic fusion parsing technology to conduct a comprehensive review of the configuration files in the image, focusing on two core areas: user permission configuration and security policy configuration. In terms of user permission review, strictly follow the principle of least privilege. By carefully analyzing the correspondence between user accounts and permissions, it evaluates whether the permissions of each user exceed the minimum required for their work. For example, by carefully checking files such as / etc / passwd and / etc / shadow in the Linux system, it accurately identifies user accounts with excessive permissions (such as unnecessary root permissions or administrator permissions), as well as situations where the password settings are too simple or are empty passwords, ensuring that security risks are eliminated at the source in user permission configuration. In terms of security policy review, for encryption policies, carefully check whether the strength of the encryption algorithm meets the current security standards and whether the management of encryption keys is standardized and rigorous. For audit policies, ensure that audit logs can comprehensively and accurately record key operation events, such as user logins and important file accesses, so that the source can be quickly traced in case of a security incident, providing solid guarantees for security auditing and data protection.
[0161] Embodiment 5
[0162] Based on Embodiment 1 or Embodiment 2 or Embodiment 3 or Embodiment 4, the specific process of step S4 is as follows:
[0163] S41: Establish an efficient data transfer interface between the scanning engine module and the result processing module. Adopt a communication mechanism based on message queues to enable real-time and stable data transmission. Define a standardized data format, and the scanning engine outputs result data according to a unified structure;
[0164] S42: Perform format verification on the received data to determine whether it conforms to the predefined standard format. If not, feedback error information to the scanning engine module, request resending or correction;
[0165] S43: Clean the received data to remove duplicate and redundant data records. For data records with missing values in some key fields, fill them reasonably according to the data distribution characteristics and business knowledge. For missing risk level fields in certain vulnerabilities, infer and fill them in combination with the known risk levels of similar vulnerabilities and the relevant descriptions of the vulnerabilities. For those that cannot be filled, mark them as data to be confirmed and store them separately for subsequent manual processing;
[0166] S44: Extract valuable features from the original data. For vulnerability data, extract features such as vulnerability type, affected software components, and scanning tool features for discovering vulnerabilities; for malware data, extract features such as file hash values, propagation path features, and infected system areas; for configuration defect data, extract features such as configuration file types, involved user privilege levels, and security policy categories. Combine the extracted features into a structured feature vector for subsequent algorithms to use;
[0167] S45: Divide the extracted feature vectors into major categories of vulnerabilities, malware, and configuration defects, and further subdivide them under each major category. For example, divide vulnerabilities into subcategories such as operating system vulnerabilities, application software vulnerabilities, and network protocol vulnerabilities; divide malware into subcategories such as worms, trojans, and viruses; divide configuration defects into subcategories such as improper user privilege configuration, encryption policy defects, and audit policy defects, forming a hierarchical classification system;
[0168] S46: Search for combinations of vulnerabilities and malware that appear simultaneously, and analyze whether there is a situation where malware uses specific vulnerabilities for propagation or attack; explore the associations between configuration defects and vulnerabilities, malware, including whether overly high user privilege configurations lead to a higher susceptibility to specific types of malware attacks or vulnerability exploitation; generate an association analysis report through association analysis.
[0169] After step S46, the following process is also included:
[0170] S47: Calculate basic statistics such as the number and occurrence frequency of various types of vulnerabilities, malware, and configuration defects, and create a risk rating model;
[0171] S48: For each problem discovered during scanning, extract the relevant feature information of its corresponding vulnerabilities, malware, and configuration defects, substitute it into the risk rating model for calculation to obtain the risk rating result, associate the risk rating result with the problem details, and clearly mark the risk level of each problem in the database;
[0172] S49: Configure multiple report templates in the template engine, define dynamic data filling areas in the template engine, associate them with the scanning result data and risk rating information in the database, and based on the open-source visualization chart library, convert various previously generated visualization charts into the form of pictures or HTML fragments, embed them in the corresponding positions of the scanning report, and generate targeted repair suggestions for vulnerabilities, malware, and configuration defects.
[0173] Data acquisition interface construction: Establish an efficient data transmission interface between the scanning engine module and the result processing module, and adopt the Kafka communication mechanism based on the message queue to ensure real-time and stable data transmission. For example, when the scanning engine completes the scanning of a virtual machine image, immediately push the result data to the result processing module through a pre-set Kafka topic to ensure that the data is not lost and arrives in order. Also, define a standardized data format, requiring the scanning engine to output the result data according to a unified structure, including but not limited to key information such as file path, vulnerability name, malware characteristics, configuration problem description, relevant timestamp, etc., for convenient subsequent parsing and processing.
[0174] Perform format verification on the received data to check whether it conforms to the pre-defined standard format. If a format mismatch is found, promptly feedback error information to the scanning engine module, requiring it to resend or correct. For example, verify whether the vulnerability name follows the common naming convention and whether the timestamp format is correct. Conduct an integrity check on the data content to ensure that key information is not missing. For example, for vulnerability data, confirm whether it contains necessary fields such as vulnerability numbers and affected versions; for configuration defect data, check whether it clearly indicates the configuration file path and specific configuration items where the problem lies. If data is found to be missing, try to supplement it from the cache or log of the scanning engine. If it cannot be supplemented, mark it as suspicious data and have manual intervention for verification later.
[0175] Clean the data to remove duplicate and redundant data records. For example, if the same vulnerability information appears multiple times when scanning the same file or image, only keep one record and update auxiliary information such as the relevant scan timestamp. Handle missing values. For data with missing values in some key fields, fill them reasonably according to the distribution characteristics of the data and business knowledge. For example, for the missing risk level field of some vulnerabilities, infer and fill it in combination with the known risk levels of similar vulnerabilities and the relevant descriptions of this vulnerability. For those that cannot be filled, mark them as data to be confirmed and store them separately for subsequent manual processing. Feature extraction: Extract valuable features from the original data. For example, for vulnerability data, extract features such as vulnerability type, affected software components, and scanning tools that discovered the vulnerabilities. For malware data, extract features such as file hash values, propagation path characteristics (such as whether it spreads through network sharing, email attachments, etc.), and infected system areas. For configuration defect data, extract features such as configuration file types, involved user privilege levels, and security policy categories, and combine these features into structured feature vectors for subsequent algorithms to use.
[0176] The machine learning decision tree algorithm takes the extracted feature vectors as input and classifies the result data into major categories such as vulnerabilities, malware, and configuration defects. Under each major category, further subdivide according to specific features. For example, classify vulnerabilities into subcategories such as operating system vulnerabilities, application software vulnerabilities, and network protocol vulnerabilities; classify malware into subcategories such as worms, trojans, and viruses; classify configuration defects into subcategories such as improper user privilege configuration, encryption policy defects, and audit policy defects, forming a hierarchical classification system and storing it in the database for convenient query and retrieval.
[0177] Association analysis: Use the Apriori algorithm for association rule mining to analyze the result data after classification and sorting. For example, find combinations of vulnerabilities and malware that appear simultaneously and analyze whether there is a situation where malware uses specific vulnerabilities to spread or attack; explore the associations between configuration defects and vulnerabilities, malware, such as whether overly high user privilege configurations lead to being more vulnerable to specific types of malware attacks or vulnerability exploitation. Through association analysis, identify weak links in the system, generate an association analysis report, and provide a more comprehensive perspective for risk assessment.
[0178] Summary statistics: Calculate basic statistics such as the quantity and occurrence frequency of various types of vulnerabilities, malware, and configuration defects. For example, count the number of discoveries of different types of vulnerabilities over a period of time, draw a trend chart, and observe the growth trend of vulnerabilities; count the infection ratios of different malware families and analyze which malware is more active in the enterprise network; count the distribution of configuration defects in each business system and identify weak links in configuration management. Display these summary statistics information in the form of visual charts (such as bar charts, line charts, pie charts, etc.) to facilitate security managers to quickly understand the overall situation.
[0179] Risk rating model construction and application, factor quantification and weight determination: For the type and harm degree of vulnerabilities, referring to the industry standard CVSS scoring system and combining with the sensitivity of the enterprise's own business, vulnerabilities are divided into high, medium, and low harm levels, and corresponding quantitative scores are assigned respectively. For example, high-risk vulnerabilities are 9 - 10 points, medium-risk vulnerabilities are 5 - 8 points, and low-risk vulnerabilities are 1 - 4 points.
[0180] Regarding the propagation characteristics and destruction potential of malware, considering factors such as propagation speed (higher scores are given to network worms with rapid propagation), infection range (whether it can spread across network segments and systems), and destruction consequences (whether it causes system paralysis, data loss, etc.), it is also quantified into different scores. For example, malware with high destruction potential is 8 - 10 points, medium is 5 - 7 points, and low is 1 - 4 points.
[0181] For the impact range and repair difficulty of configuration defects, quantification is carried out based on factors such as the system level involved in the configuration item (higher scores for core system configuration defects), the number of user groups (higher scores for configuration problems affecting a large number of users), and the technical complexity required for repair (higher repair difficulty for those requiring professional technicians and complex operations). For example, configuration defects with high impact and high difficulty are 7 - 10 points, medium is 4 - 6 points, and low is 1 - 3 points. Through methods such as expert evaluation and historical data backtesting, determine the weights of each dimension factor in the risk rating. For example, the weight of the vulnerability factor is 0.4, the weight of the malware factor is 0.3, and the weight of the configuration defect factor is 0.3, ensuring that the weight distribution reasonably reflects the contribution of each factor to the overall risk.
[0182] Risk assessment and level classification: For each problem discovered by scanning, extract the relevant characteristic information of its corresponding vulnerability, malware, and configuration defect, and substitute it into the risk rating model for calculation. For example, for a newly discovered buffer overflow vulnerability, which belongs to the high-risk vulnerability type and has a quantitative score of 10 points; considering the business importance of the system where it is located, determine the final score of the vulnerability factor to be 8 points (after considering the weight). If no associated malware and configuration defects are found, according to the model calculation, the total risk score of this problem is 8 points, corresponding to the high-risk level.
[0183] Associate and store the risk rating result with the problem details, and clearly mark the risk level of each problem in the database for convenient subsequent query and report generation.
[0184] Scan report generation, template engine configuration, design various report templates, including detailed report templates, summary report templates, etc., to meet the needs of different user groups and application scenarios. The detailed report template covers all relevant information of the virtual machine image and the details of the scan results, and is suitable for security professionals for in-depth analysis. The summary report template highlights key information, such as risk rating results and a summary of high-risk issues, to facilitate management's quick understanding of the overall situation.
[0185] Define dynamic data filling areas in the template engine, associate with fields such as scan result data and risk rating information in the database, and ensure that data can be accurately and quickly embedded into the corresponding positions when generating reports. For example, in the HTML template, use the placeholder {{image_name}} to represent the virtual machine image name. When generating the report, the template engine automatically obtains the actual image name from the database and replaces the placeholder.
[0186] Data visualization integration: Based on the open-source visualization chart library Echarts, convert various previously generated visualization charts (such as vulnerability trend charts, malware infection ratio pie charts, etc.) into the form of pictures or HTML fragments and embed them into the corresponding positions in the scan report. For example, in the "Security Situation Analysis" section of the detailed report, insert a vulnerability trend chart to intuitively show the changes in vulnerabilities and allow users to clearly understand the system security dynamics at a glance.
[0187] Generate repair suggestions and case references: According to different types of problems (such as vulnerabilities, malware, configuration defects), combined with industry best practice repair experience, generate targeted repair suggestions. For example, for a certain high-risk operating system vulnerability found, it is recommended to immediately download and install the security patch released by the official, and at the same time provide the patch download link and installation steps description; for malware infection, it is recommended to use a specific antivirus software for scanning and killing, and give a recommended list of antivirus software and operation guides.
[0188] Report format output: Use report generation tools (such as using Python's ReportLab library to generate PDF format, Jinja2 template engine combined with HTML to generate HTML format, etc.), and according to the format selected by the user, convert the report content filled with data by the template engine and embedded with visualization charts into the corresponding format. For example, when the user needs a PDF format report, the ReportLab library arranges the report content into a standard PDF file according to the preset layout, font, etc., to ensure good printing and reading effects; when the user selects the HTML format, the Jinja2 template engine generates a dynamic HTML page for easy viewing and interaction in the browser.
[0189] The export of result data realizes the seamless connection between the scan result data and other security management systems, work order management systems, or enterprise internal big data analysis platforms, facilitating the collaborative operation of subsequent processes. Study the data access requirements of the target system and develop an adapted data interface. For example, for a security management system that supports RESTful API access, write a data export function according to its API specification to convert the scan result data into JSON format data that meets the API requirements and send it to the target system through an HTTP request; for a work order management system, if Excel spreadsheet data is required, use the Pandas library in Python to organize the scan result data into Excel format and push it according to the import process of the work order system.
[0190] A virtual machine image static scanning system based on a hyper-converged platform for implementing a virtual machine image static scanning method based on a hyper-converged platform, including an image acquisition and preprocessing module, a resource scheduling and monitoring module, a comprehensive scanning engine module, and a result processing and report generation module. The image acquisition and preprocessing module is connected to the resource scheduling and monitoring module, the resource scheduling and monitoring module is connected to the comprehensive scanning engine module, and the comprehensive scanning engine module is connected to the result processing and report generation module.
[0191] The image acquisition and preprocessing module is used to create a consistent snapshot of the storage of the image to be scanned, obtain the static virtual machine image file based on the hyper-converged distributed nodes, use the indexing and retrieval algorithm, extract and record its metadata information, perform multiple rounds of hash calculation on the virtual machine image file, and compare and verify it with the preset standard hash value; the resource scheduling and monitoring module is used to collect the comprehensive resource data of the hyper-converged platform in real time and use the preset algorithm to allocate corresponding resource quotas for the scan task, including the optimal execution node of the scan task, the allocated number of CPU cores, the set memory capacity, and the network bandwidth limit; the comprehensive scanning engine module is used to perform multi-stage in-depth scanning, first conduct a preliminary screening with file fingerprints and feature matching, then exclude false positives through intelligent version comparison, and finally use static code analysis on key complex components to deeply dig out potential high-risk vulnerabilities, quickly identify and locate known malware by extracting and analyzing multi-dimensional features of various files in the virtual machine image, and conduct a comprehensive review of the configuration files in the image; the result processing and report generation module is used to comprehensively collect various result data generated by the scanning engine module during the scanning process, and use the preset algorithm to deeply classify, correlate, and summarize and statistically analyze various result data.
[0192] In summary, the virtual machine image static scanning method and system based on the hyper-converged platform provided by the present invention create a storage snapshot of the image to be scanned, use indexing and retrieval algorithms to obtain static virtual machine image files, extract and record their metadata information, and then perform integrity verification; collect all-round resource data of the hyper-converged platform in real time, and use preset algorithms to allocate corresponding resource quotas for scanning tasks; perform multi-stage in-depth scanning, first perform preliminary screening by file fingerprint and feature matching, then exclude misjudgments through intelligent version comparison, and use static code analysis to deeply dig potential high-risk vulnerabilities for key complex components, quickly identify and locate known malicious software, and conduct a comprehensive review of the configuration files in the image; use preset algorithms to deeply classify, correlate, and summarize various result data, achieving virtual machine image scanning in large-scale hyper-converged clusters in a short time.
[0193] The deep integration of the virtual machine image static scanning architecture design for the hyper-converged platform, covering the technical implementation of seamless docking and collaborative work of each link of image acquisition, comprehensive scanning, resource scheduling, and result feedback with the hyper-converged platform, gives full play to the distributed, elastic, and integrated characteristics of the hyper-converged platform. By optimizing the image parsing process and intelligently scheduling resources, the present invention can complete virtual machine image scanning in large-scale hyper-converged clusters in a short time, greatly shortening the waiting time. Through the construction of a multi-source fusion vulnerability knowledge base and multi-stage in-depth scanning technology, static detection of malicious software, and fine detection of configuration compliance in the comprehensive scanning engine module, malicious software and the like can be quickly and efficiently identified, effectively reducing risks. Through the resource scheduling and optimization module, intelligent prediction based on real-time resource data of the hyper-converged platform, an optimization algorithm for dynamic allocation strategies, and an exception adjustment mechanism, rapid emergency handling is achieved. The dynamic resource scheduling strategy ensures that the scanning task will not cause too much burden on the normal operation of the hyper-converged platform and maintains a good performance level. With the help of an advanced false alarm filtering mechanism, the present invention effectively reduces the number of false alarms, enabling administrators to focus more on truly important security issues.
Claims
1. A static scanning method for a virtual machine image based on a hyper-converged platform, characterized in that: The following steps are involved: S1: Create a consistent snapshot of the image storage to be scanned through the image acquisition and preprocessing module, obtain static virtual machine image files based on hyper-integrated distributed nodes, use indexing and retrieval algorithms, extract and record their metadata information, perform multiple rounds of hash calculations on the virtual machine image files, and compare and verify them with the preset standard hash values; S2: The resource scheduling and monitoring module collects all-round resource data of the hyper-converged platform in real time, and uses the preset algorithm to allocate corresponding resource quotas for scanning tasks, including the optimal execution node of the scanning task, the allocation of CPU cores, the setting of memory capacity and network bandwidth limit; S3: Performs multi-stage deep scanning through the integrated scanning engine module, first screening with file fingerprints and feature matching, then eliminating misjudgments through intelligent version comparison, and finally using static code analysis to dig deep into potential high-risk vulnerabilities for key complex components. It quickly identifies and locates known malware through multi-dimensional feature extraction and analysis of various files in the virtual machine image, and conducts a comprehensive review of the configuration files in the image. S4: The result processing and report generation module comprehensively collects various result data generated by the scanning engine module during the scanning process, and uses the preset algorithm to deeply classify, organize, correlate and summarize the various result data.
2. The method for static scanning of virtual machine images based on a hyper-converged platform according to claim 1, characterized in that: The specific process of creating a consistent snapshot of the image storage to be scanned by the image acquisition and preprocessing module in step S1 is as follows: S11: The image acquisition and preprocessing module receives an external startup instruction or activates the image acquisition and preprocessing module at a specified time node to start executing the task; S12: Generate a consistent snapshot for the target image storage when starting the collection by calling the snapshot creation interface provided by the distributed storage system; S13: Setting image index and retrieval based on hyper-converged distributed nodes, quickly locating and obtaining the target image in the massive static virtual machine images, obtaining image metadata information from the distributed file image storage snapshot storage, filtering out the specified index item storage snapshot storage, obtaining image metadata information from the snapshot storage, and using the preset hash algorithm to allocate the node ID and index item storage location; S14: Create a distributed image scanning temporary file storage to store the image to be scanned, obtain a list of images to be scanned, record metadata information of the images to be scanned, determine the format of the images to be scanned, store the images in the newly created distributed image scanning temporary file storage, and perform image integrity calculation.
3. The method for static scanning of virtual machine images based on a hyper-converged platform according to claim 2, characterized in that: The specific process of step S13 is as follows: S131: Obtain image metadata information from the distributed file image storage snapshot storage, and filter out key features including image name, creation time, and size as index items; S132: Using a preset hash algorithm to allocate node IDs and index item storage locations, each node generates a hash value as a node ID based on its own network address and hardware identification information; S133: For the index item, calculate its hash value, and map the index item to the corresponding node through the preset hash algorithm; S134: When a search request is received, multiple nodes are searched in parallel according to given query conditions to quickly locate the target image index.
4. The method for statically scanning a virtual machine image based on a hyper-converged platform according to claim 2, characterized in that: The specific process of step S14 is as follows: S141: Create a distributed image scanning temporary file storage to store the image to be scanned, obtain a list of images to be scanned, and capture key metadata including the image creation timestamp, version iteration information, image source, and last use time to formulate a scanning strategy; S142: Determine the image format to be scanned by reading the specified identifier and structure information in the image file header, wherein when the image format to be scanned is not in the QCOW2 format, when the image format is VMDK or RAW, start the corresponding lossless decompression and format conversion program to convert it into a unified QCOW2 format; S143: storing the image in the newly created distributed image scanning temporary file storage, and storing the QCOW2 image in the QCOW2 format and the processed non-QCOW2 format in the newly created distributed file storage location; S144: After the image acquisition, decompression, and conversion operations are completed and stored in the standard image cache, the integrity check of the system is started. For the QCOW2 format image, hash calculation is performed based on the image in the snapshot storage. For the non-QCOW2 format image, hash calculation is performed based on the QCOW2 format image converted from the image cache, and the calculated hash value is stored in the database as a reference value. S145: After a round of operations on the image, a hash value is calculated for the image, and the hash value of the current image is compared and verified with the pre-stored standard hash value. If the two are consistent, the image data is determined to be authentic and complete. If they are inconsistent, an alarm mechanism is triggered, and an alarm notification containing detailed information of the image identifier and the wrong hash value is sent to the operation and maintenance monitoring system.
5. The method for static scanning of virtual machine images based on a hyper-converged platform according to claim 1, characterized in that: In step S2, the resource scheduling and monitoring module collects all-round resource data of the hyper-converged platform in real time, and uses the preset algorithm to allocate corresponding resource quotas for the scanning task. The specific process is as follows: S21: Deploy Node Exporter and Prometheus servers, use Node Exporter to collect various basic resource indicators of the hyper-converged platform nodes, and then use Prometheus servers to aggregate, store and query data to provide data support for subsequent processes; S22: The Prometheus server writes query statements using a specified query language, summarizes the resource data stored in the database at a specified time node, establishes a resource abundance scoring mechanism, determines the optimal cluster node for executing the scanning task based on the resource abundance scoring result, allocates corresponding CPU computing power to each scanning task based on the node CPU resource status and the complexity of the scanning task, and sets the memory capacity and network bandwidth; S23: assigning priorities to tasks according to task sources and task timeliness, creating task queues for all submitted scanning tasks based on task priorities, and assigning optimal execution nodes to each task in the task queue. According to the task queues, unloading tasks to the assigned optimal execution nodes in turn for processing; S24: Set multi-dimensional resource monitoring thresholds to monitor resource usage during task execution in real time, and use a preset scheduling algorithm to adjust and update the task scheduling strategy in real time.
6. The method for static scanning of virtual machine images based on a hyper-converged platform according to claim 1, characterized in that: The specific process of step S3 is as follows: S31: Establish data connection with NVD and CVE databases, obtain the latest vulnerability information through regular data synchronization interfaces, clean data obtained from different sources, remove duplicate and invalid information, standardize the unified format, and extract key information including vulnerability number, vulnerability description, affected version range, and repair suggestions, integrate and store them into the local knowledge base, and build a unified data structure; S32: Set up a real-time monitoring component. When the NVD or CVE database releases an emergency security update or a major vulnerability is disclosed in the industry, the update process is immediately triggered to update the relevant information to the knowledge base within a specified time. S33: Through the initial screening of file fingerprints and feature matching, intelligent version comparison to eliminate misjudgment, and a progressive scanning method of key complex component code analysis, it comprehensively detects vulnerabilities from coarse-grained to fine-grained; S34: Build a malware feature recognition model, collect malware samples from public malware sample libraries and sample data shared by security agencies, extract multi-dimensional features from the samples, build feature vectors, and input them into the malware feature recognition model for training; S35: for various files in the virtual machine image to be scanned, the file hash value is calculated in sequence, the file header structure is parsed, and the semantic features of the code segment are extracted according to the feature extraction method in the training stage, a feature vector of the file to be detected is constructed, and the feature vector is input into the trained malware feature recognition model; S36: The malware feature recognition model outputs a malicious probability value for each file and compares it with a set probability threshold. When the probability value is greater than the threshold, the file is marked as suspected malware. S37: For suspected malware, intelligent feature matching technology based on machine learning models is used for verification. The model is trained based on typical feature patterns of known malware. The file features are matched with typical feature patterns including specified code segment sequences and file structure layouts. If the match is successful, the file is determined to be malware, and its location and type information in the image are recorded in detail to generate a malware identification report.
7. The method for static scanning of virtual machine images based on a hyper-converged platform according to claim 1, characterized in that: The specific process of step S4 is as follows: S41: Establish an efficient data transmission interface between the scanning engine module and the result processing module, adopt a communication mechanism based on a message queue to enable real-time and stable data transmission, define a standardized data format, and the scanning engine outputs result data according to a unified structure; S42: Perform format check on the received data to determine whether it conforms to a predefined standard format. If not, feedback error information to the scanning engine module to request resending or correction. S43: Clean the received data, remove duplicate and redundant data records, and reasonably fill in the missing data of some key fields according to the distribution characteristics of the data and business knowledge. For the missing risk level fields of some vulnerabilities, fill in the missing data by inference based on the known risk levels of similar vulnerabilities and the relevant descriptions of the vulnerabilities. For those that cannot be filled, mark them as pending data and store them separately for subsequent manual processing; S44: Extract valuable features from raw data. For vulnerability data, extract vulnerability type, affected software components, and scanning tool features that discover vulnerabilities; for malware data, extract file hash values, propagation path features, and infected system area features; for configuration defect data, extract configuration file types, involved user permission levels, and security policy category features, and combine the extracted features into a structured feature vector for use by subsequent algorithms; S45: The extracted feature vectors are classified into the categories of vulnerabilities, malware, and configuration defects, and each category is further subdivided, including vulnerabilities are classified into operating system vulnerabilities, application software vulnerabilities, and network protocol vulnerabilities; malware is classified into worms, Trojans, and viruses; configuration defects are classified into improper user permission configuration, encryption policy defects, and audit policy defects, forming a hierarchical classification system; S46: Look for combinations of vulnerabilities and malware that appear at the same time, and analyze whether there is malware that uses specified vulnerabilities to spread or attack; explore the relationship between configuration defects and vulnerabilities and malware, including whether excessive user rights configuration makes it more vulnerable to specified types of malware attacks or vulnerability exploits; generate correlation analysis reports through correlation analysis.
8. The method for statically scanning a virtual machine image based on a hyper-converged platform according to claim 7, characterized in that: After step S46, The process includes: S47: Calculate basic statistics on the number and frequency of various vulnerabilities, malware, and configuration defects, and create risk rating models; S48: For each problem found by the scan, extract relevant feature information of the corresponding vulnerability, malware, and configuration defect, substitute it into the risk rating model to calculate and obtain the risk rating result, associate the risk rating result with the problem details and store it, and clearly mark the risk level of each problem in the database; S49: Configure multiple report templates in the template engine, define dynamic data filling areas in the template engine, associate them with the scan result data and risk rating information in the database, and convert the previously generated various visualization charts into pictures or HTML fragments based on the open source visualization chart library, embed them into the corresponding positions of the scan report, and generate targeted repair suggestions for vulnerabilities, malware, and configuration defects.
9. A virtual machine image static scanning system based on a hyper-converged platform, used to implement a virtual machine image static scanning method based on a hyper-converged platform as described in any one of claims 1 to 8, characterized in that: It includes an image acquisition and preprocessing module, a resource scheduling and monitoring module, a comprehensive scanning engine module and a result processing and report generation module, wherein the image acquisition and preprocessing module is connected to the resource scheduling and monitoring module, the resource scheduling and monitoring module is connected to the comprehensive scanning engine module, and the comprehensive scanning engine module is connected to the result processing and report generation module; The image acquisition and preprocessing module is used to create a consistent snapshot of the image storage to be scanned, based on the hyper-integrated distributed nodes, use the index and retrieval algorithm to obtain the static virtual machine image file, extract and record its metadata information, perform multiple rounds of hash calculations on the virtual machine image file, and compare and verify it with the preset standard hash value; The resource scheduling and monitoring module is used to collect all-round resource data of the hyper-converged platform in real time, and use a preset algorithm to allocate corresponding resource quotas for scanning tasks, including the optimal execution node of the scanning task, the allocation of CPU cores, the setting of memory capacity and network bandwidth limit; The comprehensive scanning engine module is used to perform multi-stage deep scanning, firstly screening by file fingerprint and feature matching, then eliminating misjudgment by intelligent version comparison, and finally using static code analysis to dig deep into potential high-risk vulnerabilities for key complex components, quickly identifying and locating known malware by performing multi-dimensional feature extraction and analysis on various files in the virtual machine image, and conducting a comprehensive review of the configuration files in the image; The result processing and report generation module is used to comprehensively collect various types of result data generated by the scanning engine module during the scanning process, and use a preset algorithm to perform in-depth classification, correlation analysis, and summary statistics on various types of result data.
Citation Information
Patent Citations
Network intrusion detection method, device and apparatus and readable storage medium
CN111464526A
Virtualization-based kernel vulnerability patch verification method and device
CN116305133A