Docker mirror image vulnerability detection method and system based on instruction mapping

Through the Docker image vulnerability detection method based on instruction mapping, combined with the file path positioning of Dockerfile and the image layer, using seven-tuple feature representation and differentiated vulnerability detection, the problems of incomplete vulnerability detection and high false alarm rate in the existing technology are solved, and more efficient and accurate vulnerability detection is achieved.

CN120296744APending Publication Date: 2025-07-11NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357292.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing Docker image vulnerability detection tools have problems with inaccurate detection results and limited coverage in identifying vulnerabilities in applications and their dependencies, and existing methods cannot fully cover all potential risks, resulting in incomplete detection results and high false positive rates.

Method used

The Docker image vulnerability detection method based on instruction mapping is adopted. Through three parts: file discovery, software package feature extraction and vulnerability detection, combined with Dockerfile, image layer and file storage directory, the two-stage instruction mapping algorithm is used to obtain the corresponding file path of the software package, and the feature representation is used in the seven-tuple format, and the software package type is distinguished for differentiated vulnerability detection.

Benefits of technology

It achieves more comprehensive vulnerability coverage, reduces false positive rates, improves detection accuracy and reliability, and ensures the integrity and efficiency of vulnerability detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296744A_ABST
    Figure CN120296744A_ABST
Patent Text Reader

Abstract

The invention provides a Docker mirror image vulnerability detection method and system based on instruction mapping, relates to the technical field of Docker containers, and solves the problems of incomplete file object recognition, low detection efficiency and insufficient security vulnerability detection accuracy of different types of software packages in existing container mirror image vulnerability detection. The method comprises the following three steps: S1, for file discovery, associating a Docker file, a mirror layer and a file storage directory, and quickly positioning a file through an absolute path; s2, feature extraction: performing feature representation on a software package by using a seven-tuple format, and supplementing dependency items which may be missed by using a size verification mechanism, so as to prepare for subsequent detection and ensure the integrity of detection at the same time; and S3, for vulnerability detection, software package types are distinguished, different vulnerability detection operations are executed respectively, and a detection result is recorded in a fixed format and can be reused in subsequent detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Docker containers and is applied to the security analysis process of Docker images. Specifically, it relates to a method and system for detecting vulnerabilities in Docker images based on instruction mapping. Background Art

[0002] In modern software development and deployment environments, Docker container technology has been widely used due to its advantages such as high efficiency and light weight. The creation of container instances depends on specific image templates, which contain all the dependencies and environmental configurations required to run the application. However, quality issues in the images can directly affect the running state of the containers and may pose threats to the security and stability of the host. Therefore, ensuring image security is crucial for protecting the entire containerized application ecosystem.

[0003] To maintain the security of the images, users need to regularly update software packages and conduct security checks. However, during the image building process, many images contain potential security vulnerabilities due to the use of expired software packages, lack of version control, or failure to perform appropriate security checks. Existing open-source image vulnerability scanning tools on the market often focus on the analysis of image structures and formats, as well as the ability to search for and filter known vulnerabilities, but they are insufficient in identifying vulnerabilities in applications and their dependencies. The limitations of these tools result in inaccurate detection results and an inability to comprehensively cover all potential risks.

[0004] In addition, existing image vulnerability detection solutions usually rely on the above-mentioned scanning tools and thus inherit their defects. Some more advanced detection methods attempt to identify potential vulnerabilities in code written in programming languages such as Python through static code analysis, or to discover more details by constructing an inheritance relationship graph or a global relationship tree of the images. However, such methods also face problems such as one-sided detection results and limited generality, and their accuracy is also affected by the number of available samples.

[0005] Ideally, image vulnerability detection should be a systematic process that starts with scanning the image content, identifying various files therein, comparing the extracted software package features with the latest vulnerability database, and finally reporting the matching results. Ensuring the integrity and accuracy of the detection is the key challenge. If some files are missed during the file discovery stage, it will affect the integrity of subsequent analysis; while incorrect reporting will lead to false positives or false negatives, compromising the reliability of the detection.

[0006] To address these issues, a possible approach is to directly analyze the mirror storage directory. Although this method is intuitive, it is not ideal because file partitioning is complex and time-consuming. In contrast, using the Dockerfile can more efficiently obtain mirror construction information because the Dockerfile records every step in the mirror creation process, including the installed software packages and their paths. However, relying solely on the Dockerfile may overlook some indirect dependencies, making it impossible to achieve optimal detection results using either method alone.

[0007] To achieve the goal of efficiently and comprehensively discovering all mirror file objects, it is necessary to explore a new method that combines the analysis of the mirror storage directory, mirror layers, and Dockerfile. This comprehensive strategy aims to correlate the data among the three to ensure the identification of all relevant files without omission while improving detection efficiency.

[0008] To further improve detection accuracy, corresponding detection measures need to be taken according to different types of software packages. For application software packages, existing open-source source code analysis tools can be used for effective vulnerability detection. For non-application software packages, it is necessary to consider how to enhance the design of the detection mechanism based on vulnerability library matching to ensure that each detection result can be independently verified, providing a more reliable and accurate security assessment. Summary of the Invention

[0009] Based on the current situation in the background technology, the purpose of the present invention is to solve the problems of incomplete identification of file objects, low detection efficiency, and insufficient detection accuracy of security vulnerabilities in different types of software packages in existing container image vulnerability detection. Therefore, a Docker image vulnerability detection method and system based on instruction mapping are proposed. The present invention ensures the integrity of the detection by sequentially executing three parts: file discovery, software package feature extraction, and vulnerability detection, increases the ability to extract undetected dependencies, and reduces the vulnerability false positive rate, making the entire vulnerability detection process more complete and efficient.

[0010] The present invention adopts the following technical solutions to achieve the purpose:

[0011] A Docker image vulnerability detection method based on instruction mapping, comprising the following steps:

[0012] Step S1, File discovery: Correlate the Dockerfile, mirror layer, and file storage directory corresponding to the software package, and use a two-stage instruction mapping algorithm to obtain the path of the file corresponding to the software package, and locate the file corresponding to the software package through this path;

[0013] Step S2, Feature Extraction: Use the seven-tuple format to represent the features of the files corresponding to the located software package, and use the size verification mechanism to supplement the dependencies that may be missed during detection to determine all the files to be detected for vulnerability detection corresponding to the software package;

[0014] Step S3, Vulnerability Detection: After the files to be detected are determined, distinguish the types of software packages corresponding to the files to be detected, and perform different vulnerability detection operations according to different software package types, and obtain and record the vulnerability detection results of the software package.

[0015] Specifically, in Step S1, the paths of the files corresponding to the software package include two types. One is the installation path of the software package in the image layer, which represents the relative path of the corresponding file based on the image layer; the other is the storage path of the image layer on the host, which represents the absolute path prefix of the corresponding file based on the host. After splicing the installation path and the storage path, the actual storage location of the software package on the host is obtained, thus completing the positioning of the files corresponding to the software package.

[0016] Specifically, for the installation path, according to the intuitive description of the image building process in the Dockerfile, the installation situation of the software package is obtained by analyzing the content of the Dockerfile instructions, that is, after performing the operation of reverse generating the Dockerfile from the image, the installation path of the software package in the image layer is obtained; for the storage path, directly based on the overlay2 storage characteristics, after file association, the storage path of the image layer on the host is obtained.

[0017] Furthermore, the two-stage instruction mapping algorithm includes the logical mapping in the first stage and the physical mapping in the second stage; the logical mapping is used to achieve the matching between the Dockerfile instructions and the image layer, so as to obtain the relative path of the file based on the image layer; the physical mapping is used to achieve the physical association between the Dockerfile instructions and the storage directory of the image layer, so as to obtain the absolute path of the file based on the host.

[0018] Specifically, in the logical mapping stage, first determine the instruction content in the Dockerfile related to file import and environment configuration; when analyzing the file import instructions, combine the context semantics to determine the change of the directory environment, extract the corresponding relationship between the file name and the file import path, and record the keywords, file names and installed relative paths of the instructions in the first preset format;

[0019] Secondly, take the image layer size of 0 as a feature, filter all Dockerfile instructions with this feature, and only retain the build instructions that actually build the image layer. This build instruction forms a one-to-one mapping with the image layer, and record the keywords of the instructions and the image layer identifier in the second preset format.

[0020] Specifically, in the physical mapping phase, in the first step, the diffID and the chainID directory are associated. Given the known storage location of the chainID, the chainID value corresponding to each diffID of the image layer is calculated. After concatenating the calculated value with the storage address of the chainID directory, the association between the diffID and the chainID directory is completed;

[0021] In the second step, the diffID and the cacheID are associated by accessing the cache_id file in the chainID directory;

[0022] In the third step, the recorded value of the cache_id file in the chainID directory is concatenated with the storage path of the cacheID to determine the storage path of the image layer; the diffID, the chainID value, the cacheID value, and the corresponding storage path of the image layer are recorded in the third preset format;

[0023] In the fourth step, based on the relative path of the file installed in the image layer and the storage path of the image layer obtained, the relative path is concatenated with the storage path of the image layer to obtain the absolute path of the file stored on the host; and the keyword of the instruction, the file name, the diffID, the cacheID value, and the absolute path of the file are recorded in the fourth preset format, thereby completing the positioning of the file corresponding to the software package.

[0024] Furthermore, in step S2, first, the representation form of the software package is determined. Using the seven-tuple format, the type, name, version, size, image layer diffID, absolute path of the corresponding file, and the detection result of whether there are vulnerabilities of the software package are respectively represented;

[0025] Subsequently, based on the absolute path information of the file corresponding to the software package whose positioning has been completed, feature extraction is performed on each field value in the seven-tuple format of the software package; for the type of the software package, two numerical values are used to represent application software packages and non-application software packages respectively, and the type of the file corresponding to the software package is determined by analyzing Dockerfile instructions; for the name, image layer diffID, and absolute path of the corresponding file of the software package, they are all obtained according to the fourth preset format; for the size of the software package, three types of system commands, namely ls, du, and size, are used to obtain it; for the detection result of whether there are vulnerabilities in the software package, three numerical values are used to represent the states of undetected, vulnerable, and non-vulnerable respectively, where the undetected state uses the default state, and the default state is modified according to the detection result.

[0026] Preferably, after the representation form of the software package is determined, a size verification mechanism is used for integrity check. Taking each image layer as the verification unit, all software packages with the same diffID of the image layer are screened out, and the sum of the numerical values of the corresponding size fields is compared with the size of the image layer. If the comparison results are the same, it indicates that all the files to be inspected in the image layer have been extracted through the absolute path information positioning method. If the comparison results are different, it means that there are still missing files to be inspected in the image layer. The missing files to be inspected are discovered and extracted by traversing the subdirectories of the image layer, and the types of the software packages corresponding to the missing files to be inspected are marked as non-application software packages. Finally, taking each image layer as a unit, all the information obtained in the process of feature extraction and integrity check of the software package in the seven-tuple format is integrated to generate an image software list. Each item in the image software list corresponds to the seven-tuple format features of a software package, and the detection results of whether there are vulnerabilities in the software package are all marked as the default state initially.

[0027] Specifically, in step S3, the types of the software packages corresponding to the files to be inspected are divided into application software packages and non-application software packages, and different vulnerability detection operations are performed respectively.

[0028] The two different vulnerability detection operations are executed in parallel. Among them, the application software packages adopt the source code analysis method, taking the absolute path of the application software package as a parameter, performing the corresponding code check, and detecting the possible vulnerability information. The non-application software packages adopt the online query method, and the vulnerability query is carried out according to the preset fields of each item corresponding to the non-application software package in the image software list. The vulnerability query processes between different items are independent of each other to avoid the repeated marking of the CVE identifier for different non-application software packages.

[0029] After the vulnerability detection operations of the application software packages and the non-application software packages are both completed, the corresponding vulnerability detection results are recorded, so as to modify the detection results of whether there are vulnerabilities in the software packages of each item in the image software list and modify their initial default state. Taking each image layer as a unit, the information on whether there are vulnerabilities in all the software packages in the image layer is recorded for reusing the vulnerability detection results in the subsequent layer sharing.

[0030] The present invention also provides a computer system, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the foregoing Docker image vulnerability detection method based on instruction mapping.

[0031] In summary, due to the adoption of the present technical solution, the beneficial effects of the present invention are as follows:

[0032] The present invention aims to solve the problems of incomplete vulnerability detection scope and false positives in the prior art. By optimizing the detection process, the method ensures more comprehensive vulnerability coverage and reduces the occurrence of false positives, thereby improving the accuracy and reliability of detection.

[0033] To ensure the integrity of vulnerability detection, the present invention starts from the storage path of the image file and introduces a two-stage mapping algorithm in the file search process. In the logical mapping stage, it realizes reverse engineering of the image and generates an equivalent Dockerfile. By carefully analyzing the relevant file import instructions therein, the relative path information of the file is successfully extracted. This process lays a solid foundation for subsequent vulnerability analysis, enabling more accurate identification of potential security hazards.

[0034] After entering the physical mapping stage, the present invention further correlates the metadata related to image storage. This not only realizes the accurate positioning of the absolute path of the file but also enhances the understanding of the file system structure, helping to discover hidden vulnerabilities that may be overlooked. This deep mapping mechanism is of great significance for improving the depth and breadth of vulnerability detection.

[0035] In addition, the present invention adopts a seven-tuple feature representation form for software packages and introduces a size verification mechanism. This mechanism pays special attention to the extraction of undetected dependencies, ensuring that even indirectly dependent components are not left out. This method effectively improves the comprehensiveness of vulnerability detection and avoids blind spots in vulnerability detection caused by complex dependency relationships.

[0036] Regarding reducing the false positive rate of vulnerabilities, the present invention adopts a differentiated strategy. For non-application type software packages, the online query method is improved to improve the accuracy of information acquisition; for application software packages, static analysis at the source code level is adopted. This method can identify potential risks at an early development stage, reduce the cost of later maintenance, and also significantly reduce the possibility of false positives, ensuring that every vulnerability in the report is real and worthy of attention. Brief Description of the Drawings

[0037] Figure 1 It is a schematic diagram of the overall process of the method of the present invention;

[0038] Figure 2 It is a schematic diagram of the process of the file discovery step in the method of the present invention;

[0039] Figure 3 It is a schematic diagram of the instruction mapping record format in the method of the present invention;

[0040] Figure 4 It is a schematic diagram of the association process of diffID and cacheID in the method of the present invention;

[0041] Figure 5 It is a schematic diagram of the seven - tuple format of the software package features in the method of the present invention;

[0042] Figure 6 It is a schematic diagram of the vulnerability result record format in the method of the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0044] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0045] Embodiment

[0046] This embodiment will elaborate on the preferred content, related explanations, meanings, etc. of the Docker image vulnerability detection method based on instruction mapping of the present invention.

[0047] It can be referred to Figure 1 for the overall process schematic. The Docker image vulnerability detection method based on instruction mapping will include the following steps:

[0048] Step S1, File discovery: Associate the Dockerfile, image layer, and file storage directory corresponding to the software package, use a two - stage instruction mapping algorithm to obtain the path of the file corresponding to the software package, and locate the file corresponding to the software package through this path;

[0049] Step S2, Feature extraction: Use the seven - tuple format to represent the features of the file corresponding to the located software package, use a size verification mechanism to supplement the dependencies with the possibility of missed detection, and determine all the files to be detected for vulnerability detection corresponding to the software package;

[0050] Step S3, Vulnerability detection: After the files to be detected are determined, distinguish the types of software packages corresponding to the files to be detected, and perform different vulnerability detection operations according to different software package types, and obtain and record the vulnerability detection results of the software package.

[0051] The detailed content of the file discovery part of Step S1 is introduced as follows.

[0052] In the field of Docker container technology, based on the characteristics of the overlay2 file storage driver, the merged layer provides a unified file view after union mounting for Docker users. However, the merged layer does not contain any file entities, and the real file entities exist in the lowerdir and upperdir layers. When writing to an image file, the copy-on-write feature only allows overlay2 to modify the file copies in the container layer; when deleting an image file, overlay2 creates a whiteout file in the container layer to prevent the display of the image layer files. That is to say, any operation on the image file will only be carried out in the container layer, and the image layer remains unchanged. If files are obtained from the perspective of the merged layer, it will lead to the omission of some image layer files, affecting the integrity of the detection scope. Therefore, it is necessary to extract and analyze the files in the image layer, which is the main content and purpose of step S1.

[0053] In the Linux operating system, the method of installing software packages through the package manager has the characteristic of a fixed storage path. Users can use a directed query operation with a time complexity overhead of O(1) to conveniently obtain the list of software packages therein. Compared with the directed query of the fixed path, the method of obtaining software package information by traversing all the contents of the image layer has a relatively large time overhead. Therefore, referring to the characteristics of the package manager in the Linux system, by locating the storage paths of each software package in the image, the efficiency of software package search and feature extraction can be improved.

[0054] In this embodiment, in order to achieve the location of the software package, two paths need to be obtained. One is the installation path of the software package in the image layer, and the other is the storage path of the image layer on the host. By simply splicing the two paths, the real storage location of the software package on the host can be obtained. For path one, considering that the Dockerfile can intuitively describe the image building process, the installation situation of the software package can be obtained by analyzing the instruction content thereof, that is, after performing the operation of reverse generating the Dockerfile from the image, the installation path of the software package in the image layer can be obtained. For path two, it can be obtained based on the file association of the overlay2 storage characteristics.

[0055] The Docker daemon supports clients to view the history of image building through relevant commands (such as docker history). This history contains the time, method, instructions, and size of each image layer during building. By extracting the instruction information from it, the corresponding Dockerfile of the image can be restored. However, the number of instructions in the Dockerfile is often more than the number of image layers. Many instructions related to environment configuration do not create additional image layers but only add corresponding information to the metadata of the image, resulting in the inability to directly match the instructions one-to-one with the image layers.

[0056] In this embodiment, to solve the problem of matching Dockerfile instructions with image layers in the process of obtaining Path 1 and the file association in the process of Path 2, a two-stage instruction mapping algorithm is adopted, which includes the logical mapping in the first stage and the physical mapping in the second stage. The logical mapping in the first stage is used to achieve the matching of Dockerfile instructions with image layers and obtain the relative path of the file based on the image layer. The physical mapping in the second stage is used to achieve the physical association between the Dockerfile instructions and the storage directory of the image layer and obtain the absolute path of the file based on the host. The content of this part can be seen in Figure 2 the process schematic diagram of.

[0057] In the logical mapping stage, first, focus on the instruction content in the Dockerfile related to file import and environment configuration. Each layer of the image in Docker has an independent Rootfs, and all content is imported based on the root path. During this process, Dockerfile keywords and nested shell instructions may change the file path, such as the ENV keyword and the cd instruction. When analyzing the file import instructions, it is necessary to timely discover the change of the directory environment in combination with the context semantics, aiming to extract the correct corresponding relationship between the file name and the file import path. Subsequently, according to Figure 3 the format (a) shown in, record the instructions, file names, and installed relative paths.

[0058] Secondly, filter the Dockerfile instructions. Many instructions related to environment configuration do not create corresponding image layers after being parsed by the Docker engine but are recorded in the config file in JSON format. These instructions will still be recorded in the image building history provided by the Docker daemon, but they have the characteristic that the size of the image layer is 0. Therefore, in this embodiment, all instructions are filtered according to this characteristic, and only the instructions that actually build the image layer are retained. The instructions will achieve a one-to-one mapping with the image layer, and at the same time, according to Figure 3 the format (b) shown in, record the keyword instructions and the image layer identifier.

[0059] Any layer of the image has three IDs, namely diffID, chainID, and cacheID. The diffID is the sha256 value calculated based on the layer content and serves as the unique identifier of the layer. The chainID and cacheID are directory names related to the storage of the image layer. The data of the image layer is stored under cacheID, and the chainID provides an indexing function for cacheID. Users can calculate the corresponding content-addressable index chainID value based on the existing information and then associate it with the cacheID of the image layer.

[0060] The specific method for calculating the chainID is shown in the following formula:

[0061] chainID i = SHA256(chainID i-1 + diffID i )

[0062] In the formula, i ≥ 1, and i represents the sequential number of the image layer from bottom to top. When i = 1, chainID1 = diffID1, and the SHA256(…) function is used to generate a 256-bit hash value.

[0063] In the physical mapping stage, the first step is to associate the diffID and the chainID directory. The storage location of the chainID is known and is located in the following directory:

[0064] / var / lib / docker / image / overlay2 / layerdb / sha256.

[0065] By using the method for calculating the chainID value, the chainID value corresponding to the diffID of each image layer is obtained. By concatenating the result value and the storage address of the chainID directory, the association between the diffID and the chainID directory can be achieved. The method for calculating the chainID directory path is shown in the following formula:

[0066]

[0067] In the formula, Path chainID represents the chainID directory path.

[0068] In the second step, associate the diffID and cacheID directories. The chainID directory contains multiple files. For example, the file named cache_id records the cacheID value of the directory where the mirror data is actually stored, the file named diff records the diffID value corresponding to the current chainID directory, and the directory named parent records the diffID value of the parent mirror layer of the current layer. By accessing the cache_id file in the chainID directory, the association between the diffID and cacheID can be achieved. Figure 4 Shows the association process between the diffID and cacheID.

[0069] In the third step, determine the storage path of the mirror layer. The storage location of the cacheID is in the / var / lib / docker / overlay2 directory. Concatenate the value recorded in the cache_id file obtained from the chainID directory with the storage path of the cacheID to determine the storage location of the mirror layer data. At the same time, record the diffID, chainID value, cacheID value, and mirror layer storage path according to the format (c) shown in Figure 3 .

[0070] The specific method for calculating the mirror layer storage path is shown in the following formula:

[0071]

[0072] Finally, determine the absolute path of the file in the mirror. The above steps sequentially obtain the relative path where the file is installed in the mirror layer Rootfs and the storage path of the mirror layer. Concatenate the relative path and the mirror layer storage path to obtain the absolute path where the file is stored on the host. By performing a combined query on the three records in the formats (a), (b), and (c) in Figure 3 , extract the relative path field in the format (a) and the cacheID host storage path field in the format (c) in Figure 3 , and calculate to obtain the absolute path of the file. Record the keyword instruction, file name, diffID, cacheID value, and file absolute path according to the format (d) in Figure 3 .

[0073] The specific method for calculating the absolute storage path of the file is shown in the following formula:

[0074]

[0075] In the formula, absolut_path represents the absolute path of the file, and relative_path represents the relative path of the file. Thus, the positioning process of the file corresponding to the software package through the path is realized.

[0076] The detailed content of the feature extraction part in step S2 is introduced as follows.

[0077] Locating files by absolute path can improve the search efficiency, but there may be some files missed in this way. For example, when installing some software packages, the system will automatically download the required dependencies, and these dependencies may not be shown in the Dockerfile content, resulting in incomplete detection. Although this problem does not often occur in the Dockerfile of the official Docker Hub image, it cannot represent all situations, and the check of dependencies during the process still cannot be ignored. In this embodiment, by extracting the seven-tuple features of the software package and using the size feature for verification on this basis, it aims to balance the integrity and efficiency of detection.

[0078] First, determine the software package representation form. In addition to the name and version number, other information of the software package needs to be considered to meet the vulnerability detection process. There are three types of software packages in the image, and different software packages will adopt different vulnerability detection methods, so the types of software packages need to be distinguished. To ensure the integrity of file extraction, the designed verification mechanism will be executed based on the file size feature, and fields need to be added to record the file size. There may be cases where the software names are the same but the versions are different in different image layers, so the image layers need to be distinguished. After detecting vulnerabilities, considering the reusability of the detection results, fields need to be added to represent the detection results of the software package. If further analysis of the software package content is to be carried out, the absolute path needs to be used for location.

[0079] Therefore, this embodiment uses a seven-tuple format to represent the type, name, version, size, image layer diffID, absolute path of the corresponding file, and the detection result of whether there is a vulnerability of the software package respectively. Here, you can refer to Figure 5 the brief schematic illustration.

[0080] Subsequently, extract the software package feature values. Use the recorded absolute path information to quickly locate the software package, and extract the values of each field in the seven-tuple in turn. The detailed introduction of each field is as follows:

[0081] For the type of software package, two numerical values are used to represent the corresponding application software package and non-application software package (i.e., system package and dependency). The type identification bit corresponding to the file can be determined by analyzing the Dockerfile instructions; for example, the packages imported by the user using the ADD and COPY keywords will be marked as application packages, and the packages installed using apt-get in the Debian system will be recognized as non-application software packages.

[0082] For the name, image layer diffID, and absolute path of the software package, they can all be based on Figure 3Obtained from the record in format (d). For the version field value, if the software version information is not included in the Dockerfile instruction, built-in commands such as version can be used for query. It should be noted that the application software package may sometimes be code developed independently by the user, and the version value is specified by the user, which has no reference value, and the field value can be set to be empty.

[0083] For the size field value, system commands such as ls, du, and size can be used to obtain it. For the detection result of whether there are vulnerabilities in the software package, three numerical values are used to represent the states of undetected, having vulnerabilities, and not having vulnerabilities respectively. Among them, the undetected state uses the default state, and the default state is modified according to the detection result.

[0084] As the preference for this part of this embodiment, a size verification mechanism is used for integrity check. Taking each image layer as the verification unit, filter the seven-tuple records with the same diffID of all image layers, and compare the sum of the size field values of them with the image layer size. If the two values are equal, it means that the files to be detected in the image have been extracted by the method of locating files through the absolute path. If the values are not equal, it means that there are still missing files in the image layer. Analysis finds that these files are all necessary dependencies automatically installed by the package manager during the image building process, and it is necessary to discover and extract these dependencies by traversing the subdirectories of the image layer. Their type identification bits are all marked as non-application software packages, and the name of the software package, the image layer diffID, and the absolute path will be recorded according to the current environment. The settings and acquisitions of the three fields of version, size, and vulnerability identification are consistent with the foregoing content of this step in this embodiment.

[0085] Finally, an image software list is generated. Taking the image layer as the unit respectively, integrate the information obtained in the two processes of software package feature value extraction and integrity check, and the software list of the image can be obtained. Each list item corresponds to a software package seven-tuple record, and the vulnerability identification bit is in the default state.

[0086] The detailed content of the vulnerability detection part in step S3 is introduced as follows.

[0087] Among the usage situations of different types of software packages, system software packages and dependencies, as necessary options for providing the basic running environment, often have higher popularity. The organizations to which they belong will promptly repair the disclosed vulnerabilities and update the versions regularly, and the public vulnerability databases also record the vulnerability situations of different versions of these software. Application software packages are more inclined to implement specific functional requirements, and are often application programs developed by the software owner. The vulnerabilities existing in them do not provide relevant information for users.

[0088] For this feature, vulnerability detection operations can be performed in different ways according to the classification of non-application software packages and application software packages. Specifically, in implementation, the CVE vulnerability query method is usually adopted for the detection of non-application software packages, and the source code analysis method is adopted for application software packages; in this embodiment, parallel processing is first performed between the two methods.

[0089] In terms of the detection of non-application software packages, in the detection tool architecture represented by Clair, the CVE vulnerability database module regularly obtains vulnerability metadata from the configured public source for storage, and obtains the vulnerability situation of a specific software package by querying the database content. The time for different tools to regularly update the local database ranges from several hours to several days. If the update time is too long, the incomplete list of known vulnerabilities in the local database may cause the tool to lack the ability to detect some vulnerabilities. If the update is too frequent, new network overhead problems will arise. In addition, during the processing of detection results, if a vulnerability is found in a certain software package, some tools will mark all its related dependencies with the same CVE identifier. Therefore, the main problems of existing tools are reflected in two aspects: low vulnerability detection rate and high false alarm rate. This is the main problem that this step of this embodiment wants to solve for non-application software packages.

[0090] This embodiment improves the existing tool scanning method. First, in order to avoid the problem of the update interval of the local vulnerability database and reduce the storage overhead of the local database, this embodiment changes the vulnerability query method from querying the local vulnerability database to online query, and can expand the setting of the public vulnerability source query list. Secondly, each item in the list is queried for vulnerabilities based on the file name and version information, and the query requests for different items in the list are relatively independent, so as to avoid the repeated marking of different software packages by the CVE identifier.

[0091] In this embodiment, for the vulnerability detection of application software packages, the Sonar Scanner tool is used to perform source code analysis in the local environment. Sonar Scanner is a scanning tool for the SonarQube static code analysis platform, which supports many programming languages such as Java, JavaScript, and Python, and can perform corresponding code checks according to the characteristics and rules of different languages. Taking the absolute path of the application software package as a parameter, Sonar Scanner scans the source code in the corresponding directory to detect possible vulnerability information.

[0092] Finally, record the detection results of the software packages. Modify the vulnerability identification bit in the mirror software item list to change its initial default state. Taking the mirror layer as the unit, record the vulnerability information existing in all software packages in the layer to facilitate the reuse of detection results in subsequent layer sharing. For an example of recording the detection results, please refer to Figure 6 the format.

[0093] The above method steps of this embodiment can be applied to a computer system, which includes a memory, a processor, and a computer program stored on the memory. By executing the computer program through the processor, the steps of the Docker image vulnerability detection method based on instruction mapping of this embodiment can be implemented.

Claims

1. A Docker image vulnerability detection method based on instruction mapping, characterized in that The method includes the following steps: Step S1, File discovery: Associate the Dockerfile, image layer, and file storage directory corresponding to the software package, use a two-stage instruction mapping algorithm to obtain the path of the file corresponding to the software package, and locate the file corresponding to the software package through this path; Step S2, Feature extraction: Use a seven-tuple format to represent the features of the file corresponding to the software package that has been located, use a size verification mechanism to supplement the dependencies that may be missed in detection, and determine all the files to be detected that need to be subject to vulnerability detection for the software package; Step S3, Vulnerability detection: After the files to be detected are determined, distinguish the types of software packages corresponding to the files to be detected, and perform different vulnerability detection operations according to different software package types, and obtain and record the vulnerability detection results of the software package.

2. The Docker image vulnerability detection method according to claim 1, wherein: In Step S1, the paths of the files corresponding to the software package include two types. One is the installation path of the software package in the image layer, which represents the relative path of the corresponding file based on the image layer; the other is the storage path of the image layer on the host, which represents the absolute path prefix of the corresponding file based on the host. After splicing the installation path and the storage path, the actual storage location of the software package on the host is obtained, thereby completing the location of the file corresponding to the software package.

3. The Docker image vulnerability detection method according to claim 2, wherein: For the installation path, according to the intuitive description of the image building process in the Dockerfile, obtain the installation situation of the software package by analyzing the content of the Dockerfile instructions, that is, after performing the operation of reverse generating the Dockerfile from the image, obtain the installation path of the software package in the image layer; For the storage path, directly based on the overlay2 storage characteristics, obtain the storage path of the image layer on the host after file association.

4. The Docker image vulnerability detection method according to claim 3, wherein: The two-stage instruction mapping algorithm includes logical mapping in the first stage and physical mapping in the second stage; logical mapping is used to achieve the matching between Dockerfile instructions and the image layer, so as to obtain the relative path of the file based on the image layer; physical mapping is used to achieve the physical association between Dockerfile instructions and the storage directory of the image layer, so as to obtain the absolute path of the file based on the host.

5. The Docker image vulnerability detection method according to claim 4, wherein: In the logical mapping stage, first determine the instruction content in the Dockerfile related to file import and environment configuration; when analyzing the file import instruction, combine the context semantics to determine the change of the directory environment, extract the corresponding relationship between the file name and the file import path, and record the keyword, file name, and installed relative path of the instruction according to the first preset format; Secondly, use the feature that the size of the image layer is 0 to filter all Dockerfile instructions with this feature, and only retain the build instructions that actually build the image layer. This build instruction forms a one-to-one mapping with the image layer, and record the keyword of the instruction and the image layer identifier according to the second preset format.

6. The Docker image vulnerability detection method according to claim 4, characterized in that: In the physical mapping stage, in the first step, the diffID and the chainID directory are associated. Given the known storage location of the chainID, the chainID value corresponding to each diffID of the image layer is calculated. After concatenating the calculated value with the storage address of the chainID directory, the association between the diffID and the chainID directory is completed; In the second step, the diffID and the cacheID are associated by accessing the cache_id file in the chainID directory; In the third step, the recorded value of the cache_id file in the chainID directory is concatenated with the storage path of the cacheID to determine the storage path of the image layer; the diffID, the chainID value, the cacheID value, and the corresponding storage path of the image layer are recorded in accordance with the third preset format; In the fourth step, based on the relative path of the obtained file installed in the image layer and the storage path of the image layer, the relative path is concatenated with the storage path of the image layer to obtain the absolute path of the file stored on the host; the keyword of the instruction, the file name, the diffID, the cacheID value, and the absolute path of the file are recorded in accordance with the fourth preset format, thereby completing the positioning of the file corresponding to the software package.

7. The Docker image vulnerability detection method according to claim 1, wherein: In step S2, first, the representation form of the software package is determined. Using the seven-tuple format, the type, name, version, size, diffID of the image layer, the absolute path of the corresponding file, and the detection result of whether there are vulnerabilities of the software package are respectively represented; Subsequently, based on the absolute path information of the file corresponding to the software package whose positioning has been completed, feature extraction is performed on each field value in the seven-tuple format of the software package; for the type of the software package, two numerical values are used to represent the application software package and the non-application software package respectively, and the type of the file corresponding to the software package is determined by analyzing the Dockerfile instruction; for the name, diffID of the image layer, and the absolute path of the corresponding file of the software package, they are all obtained based on the absolute path information; for the size of the software package, three types of system commands, namely ls, du, and size, are used to obtain it; for the detection result of whether there are vulnerabilities in the software package, three numerical values are used to represent the states of undetected, vulnerable, and non-vulnerable respectively, where the undetected state uses the default state, and the default state is modified according to the detected result.

8. The Docker image vulnerability detection method according to claim 7, wherein: After the representation form of the software package is determined, a size verification mechanism is used for integrity check. Taking each image layer as the verification unit, all software packages with the same diffID of the image layer are screened out, and the sum of the numerical values of the corresponding size fields is compared with the size of the image layer; if the comparison results are the same, it means that all the files to be inspected in the image layer have been extracted through the absolute path information positioning method; If the comparison results are different, it indicates that there are still missing files to be inspected in the mirror layer. The missing files to be inspected are discovered and extracted by traversing the subdirectories of the mirror layer, and the type of the software package corresponding to the missing file to be inspected is marked as a non-application software package. Finally, taking each mirror layer as a unit, all the information obtained from the feature extraction and integrity check process of the software package in the seven-tuple format is integrated to generate a mirror software list. Each item in the mirror software list corresponds to the seven-tuple format features of a software package, and the detection results of whether there are vulnerabilities in the software package are initially marked as the default state.

9. The Docker image vulnerability detection method according to claim 1, wherein: In step S3, the types of the software packages corresponding to the files to be inspected are distinguished into application software packages and non-application software packages, and different methods of vulnerability detection operations are performed respectively. The two different methods of vulnerability detection operations are executed in parallel. Among them, for the application software package, the source code analysis method is adopted, and the absolute path of the application software package is used as a parameter to execute the corresponding code check to detect the possible vulnerability information. For the non-application software package, the online query method is adopted to query for vulnerabilities according to the preset fields of each item corresponding to the non-application software package in the mirror software list. The vulnerability query processes between different items are independent of each other to avoid the repeated marking of the CVE identifier for different non-application software packages. After the vulnerability detection operations for both the application software package and the non-application software package are completed, the corresponding vulnerability detection results are recorded, so as to modify the detection results of whether there are vulnerabilities in the software packages of each item in the mirror software list and modify their initial default state. Taking each mirror layer as a unit, the information on whether there are vulnerabilities in all the software packages in the mirror layer is recorded for reusing the vulnerability detection results in the subsequent layer sharing.

10. A computer system, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the Docker image vulnerability detection method according to any one of claims 1-9.