A Risk Prediction Method for Container Image Repositories Based on Knowledge Graph
By constructing a knowledge graph-based risk prediction method for container mirror warehouses, the problem of difficulty in determining the mirror range in the prior art is solved, and the rapid and accurate security risk assessment and identification of container mirror warehouses is achieved.
Patent Information
- Application Number
- CN202310594277.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-05-23
AI Technical Summary
The prior art lacks an overall security prediction method and a security threat warning mechanism for container mirror warehouses, making it difficult to quickly and accurately determine the affected mirror range.
Build a container mirror warehouse risk prediction method based on knowledge graph, and quantify the mirror range and vulnerability range by building the mirror layer, software layer and vulnerability layer, and output risk prediction results in combination with the image download volume.
Provides a more complete and accurate overview of security situations, with better scalability, can quickly identify potential security risks, reduce performance overhead, and improve the overall security protection level of mirror warehouses.
Smart Images

Figure CN116681131B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of knowledge graph, container virtualization technology, and network security technology, and particularly relates to a method for predicting risks of a container image repository based on a knowledge graph. Background Art
[0002] A knowledge graph is a data structure that organizes and represents knowledge in a structured manner, using nodes and edges to represent the relationships between entities. The knowledge graph can process, analyze, and interpret data more effectively and comprehensively, and helps to assist decision-making, automation, and the construction of AI-based applications.
[0003] Container-based Virtualization has been widely used in software development and operation, cloud computing, microservices, etc. in recent years due to its lightweight and standardized characteristics. However, in recent years, network security incidents related to container technology have occurred frequently, which has attracted people's attention to the security of container technology. Docker is the current de facto industry standard for container technology. It encapsulates software programs and their running environments into images in a standardized format for easy storage and circulation. When in use, the container runtime parses the image content and instantiates it into a container. Images are usually hosted in a registry, and users can directly use the public images in the registry or perform secondary encapsulation on this basis. If there are security vulnerabilities in the image, then during the process of propagation and use, this security risk will be spread and amplified. Therefore, ensuring the security of container images has become an important aspect of container security.
[0004] For container image repositories, there is currently a lack of an overall security prediction method and a security threat warning mechanism, and the overall security risks of current container image repositories are unknown. When a security incident occurs, it is difficult to quickly and accurately determine the scope of affected images. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for predicting risks of a container image repository based on a knowledge graph, which helps to improve the overall security protection level of the container image repository and effectively reduce the security risks of container technology in practice.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A method for predicting risks of a container image repository based on a knowledge graph, the method for predicting risks of a container image repository based on a knowledge graph includes:
[0008] Step 1, construct a basic container image knowledge graph;
[0009] Step 1-1: Obtain the original image data, extract the image entities, attributes, and relationships of the original image data to obtain valid images, and form the image layer of the container image knowledge graph.
[0010] Step 1-2: Use dynamic analysis methods to parse the valid images, construct a software bill of materials, and use the parsed software as the software layer of the container image knowledge graph.
[0011] Step 2: Supplement software knowledge and vulnerability knowledge to complete the construction of the final container image knowledge graph.
[0012] Step 2-1: Parse the online software repository to obtain software knowledge and supplement it to the software layer.
[0013] Step 2-2: Obtain specific open-source projects, compile and extract software knowledge to supplement it to the software layer.
[0014] Step 2-3: Parse the online vulnerability library to obtain vulnerability knowledge. After matching the vulnerability knowledge with the software parsed in the software bill of materials, form the vulnerability layer of the container image knowledge graph.
[0015] Step 2-4: Receive regularly input security vulnerabilities as vulnerability knowledge and supplement it to the vulnerability layer.
[0016] Step 3: Based on the image layer, software layer, and vulnerability layer of the container image knowledge graph, quantify the image scope to obtain the set of vulnerabilities corresponding to the images in the container image repository to be predicted, and quantify the vulnerability scope to obtain the set of images affected by specific vulnerabilities in the container image repository to be predicted.
[0017] Step 4: Output the risk prediction result of the container image repository to be predicted according to the set of vulnerabilities, the set of images, and the image download volume.
[0018] The following also provides several optional methods, which are not additional limitations to the above overall solution, but only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution alone, or multiple optional methods can be combined with each other.
[0019] Preferably, the obtaining of the valid images includes:
[0020] Perform knowledge fusion and disambiguation on the extracted image entities, attributes, and relationships, and discard images that do not contain the SHA256 attribute to obtain valid images.
[0021] Preferably, the using of dynamic analysis methods to parse the valid images and construct a software bill of materials includes:
[0022] Load the container running environment.
[0023] Run the container image in a container runtime environment, start a dynamic analysis tool to monitor the behavior of the container image, obtain a list of all currently installed software components of the container image, and identify the software components loaded and executed by the container image;
[0024] Determine the dependencies existing between the software for each software component, and establish a software composition model of the container;
[0025] When the dynamic analysis by the dynamic analysis tool is completed, output a software bill of materials, which includes the names, versions of the software components used by the container image, and the dependencies between the software.
[0026] Preferably, quantifying the image scope to obtain a set of vulnerabilities corresponding to the images in the container image repository to be predicted includes:
[0027] Obtain the image, and find all directly dependent package entities according to the direct dependency relationship between the image and the software;
[0028] For each package entity, calculate the linear ordering of the package entity according to the dependency relationship of the package entity in the container image knowledge graph;
[0029] After obtaining the linear ordering, calculate the shortest path between two package entities;
[0030] Find all connected vulnerability entities in the shortest path along the shortest path;
[0031] Collect all the discovered vulnerability entities to obtain the final set of vulnerabilities corresponding to the image.
[0032] Preferably, quantifying the vulnerability scope to obtain a set of images affected by a specific vulnerability in the container image repository to be predicted includes:
[0033] Obtain the vulnerability, and find all directly dependent package entities according to the direct dependency relationship between the vulnerability and the software;
[0034] For each package entity, calculate the reverse linear ordering of the package entity according to the dependency relationship of the package entity in the container image knowledge graph;
[0035] After obtaining the reverse linear ordering, calculate the shortest path between two package entities;
[0036] Find all connected image entities in the shortest path along the shortest path;
[0037] Collect all the discovered image entities to obtain the final set of images affected by the vulnerability.
[0038] Preferably, outputting a risk prediction result of the container image repository to be predicted according to the set of vulnerabilities, the set of images, and the image download volume includes:
[0039] The risk prediction result of the container image repository to be predicted is obtained based on the image security scores of each image in the container image repository, and the image security score is obtained based on the vulnerability set, the image set, and the image download volume.
[0040] Preferably, the calculation formula of the image security score is as follows:
[0041] S T = α * S I - β * S P - γ * S V
[0042]
[0043] In the formula, S T is the image score, S I is the image download volume security score, α is the weight of S I , S P is the software security score, β is the weight of S P , S V is the vulnerability threat score, γ is the weight of S V , S SUM is the image security score, S Tmin is the minimum image score in the container image repository to be predicted, S Tmax is the maximum image score in the container image repository to be predicted.
[0044] Preferably, the calculation formula of the image download volume security score is as follows:
[0045]
[0046] In the formula, I P is the image download volume, I Pmax is the maximum download volume of the images in the container image repository to be predicted.
[0047] Preferably, the calculation process of the software security score is as follows:
[0048] Based on the vulnerability set, obtain the total number of software in the image;
[0049] Based on the image set, obtain the number of software affected by vulnerabilities in the image;
[0050] According to the number of software affected by vulnerabilities and the total number of software, obtain the proportion of affected software;
[0051] Calculate the software security score as follows:
[0052]
[0053] Wherein, P V is the proportion of affected software in the mirror, and P Vmin is the minimum proportion of affected software in the images in the container image repository to be predicted, and P Vmax is the maximum proportion of affected software in the images in the container image repository to be predicted.
[0054] Preferably, the vulnerability threat score calculation formula is as follows:
[0055]
[0056]
[0057] Wherein, V CS is the weighted sum vulnerability score of the vulnerability set corresponding to the image, n h is the number of high-risk vulnerabilities in the vulnerability set corresponding to the image, and n l is the number of low-risk vulnerabilities in the vulnerability set corresponding to the image, is the score of the i-th high-risk vulnerability, is the score of the j-th low-risk vulnerability, a is the weight of high-risk vulnerabilities, b is the weight of low-risk vulnerabilities, and S V is the vulnerability threat score, is the minimum weighted sum vulnerability score in the container image repository to be predicted, is the maximum weighted sum vulnerability score in the container image repository to be predicted.
[0058] The present invention provides a method for predicting risks in a container image repository based on a knowledge graph. Compared with the prior art, it has the following beneficial effects:
[0059] (1) Performing security prediction on container images based on a knowledge graph can capture the complex relationships between different entities and provide a more complete and accurate overview of the overall security situation.
[0060] (2) It has better scalability. The knowledge graph can be flexibly expanded, which is beneficial for predicting a large number of container images and identifying potential security risks.
[0061] (3) Through the quantification of the security threat scope, it is possible to efficiently find out the vulnerability situation contained in the image and the scope of images affected by specific vulnerabilities, and have a faster and more accurate understanding of the security situation of the image and the scope of vulnerability hazards. Compared with a comprehensive security scan of the image repository, this method can greatly reduce the performance overhead.
[0062] (4) By designing a security assessment algorithm, comprehensively considering the download volume, number of vulnerabilities, vulnerability scores, etc. of the images, more accurately assessing the overall security of the image repository, and screening the distribution of high-risk images based on the security assessment to help protect the company and individuals from potential security vulnerabilities and other risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flowchart of a method for predicting risks in a container image repository based on a knowledge graph according to the present invention;
[0064] Figure 2 It is a flowchart for constructing a container image security knowledge graph according to the present invention;
[0065] Figure 3 It is a schematic diagram of quantifying and searching the image scope according to the present invention;
[0066] Figure 4 It is a schematic diagram of quantifying and searching the vulnerability scope according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0069] The present invention proposes a method for assessing security risks in a container image repository based on a knowledge graph, representing container images and the version information of included software packages in a structured manner, and maintaining the mutual dependencies between container images. By parsing the dependency relationships, the scope of influence of security risks can be quickly located in a less costly manner without having to rescan the image content. The present invention defines different threat levels for images, and based on the proportion of high-risk images, the overall security of the container image repository can be effectively evaluated. A semi-artificial security data supplementation method based on a search engine is proposed to address the problem of lagging security data updates. By regularly and automatically retrieving the latest security events and combining manual assistance to extract information, the timeliness and effectiveness of the data in the knowledge graph can be ensured. The present invention helps to improve the overall security protection level of the container image repository and effectively reduce the security risks in the practice of container technology.
[0070] As Figure 1As shown in the figure, a method for predicting risks of a container image repository based on a knowledge graph proposed in this embodiment includes the following steps:
[0071] Step 1. Construct a basic container image knowledge graph: Collect original image data from the public image repository Docker Hub, parse and extract entities and relationships, parse the software bill of materials (SBOM) of the image through dynamic analysis methods. For software components that cannot be recognized, use binary feature matching methods for recognition, and finally organize and combine them into image component dependency relationships and vulnerability knowledge. As Figure 2 shown, the specific process is as follows.
[0072] Step 1-1. Obtain original image data, extract the image entities of the original image data and the attributes and relationships of the image entities to obtain valid images, and form the image layer of the container image knowledge graph.
[0073] In this embodiment, original image data is collected from the open-source image repository Docker Hub by means of web crawlers, etc. The images are collected according to the download volume as an indicator (for example, only collect images with a download volume higher than 10,000), and the image entities and their attributes and relationships are extracted to obtain structured or semi-structured data.
[0074] In addition, after extracting the structured or semi-structured data, perform knowledge fusion and disambiguation to form a unified knowledge representation, and discard the image knowledge that does not contain the SHA256 attribute, which means deprecating the image to ensure the quality of the knowledge graph library, and thus form the image layer of the knowledge graph.
[0075] Step 1-2. Use dynamic analysis methods to parse the valid images, construct a software bill of materials, and use the parsed software as the software layer of the container image knowledge graph.
[0076] In this embodiment, dynamic analysis methods are used to parse valid images (Docker images in this embodiment), construct a software bill of materials (SBOM), which involves running the container image in a controlled environment and observing its behavior to determine its software composition. This method can provide more accurate and comprehensive software component information contained in the container image, especially including some software that may be dynamically loaded at runtime.
[0077] More specifically, the steps for dynamically analyzing and constructing the SBOM are as follows:
[0078] a) Load a container running environment Docker to support running the container image.
[0079] b) Run the container image in the operating environment and start the dynamic analysis tool to monitor the behavior of the container. The dynamic analysis tool used in this embodiment is the apt package manager, and a list of all currently installed software components of the container image is obtained through the command "apt list --installed".
[0080] c) As the container image runs, the dynamic analysis tool will continuously observe its behavior and identify the software components loaded and executed by the container image. In this embodiment, software components, software packages, and packages all refer to a software package that constitutes the image, and different descriptions are used in different contexts.
[0081] d) The tool collects information about the packages, libraries, and other components used by the container image, parses the data to identify the software components used by the container, and constructs a dependency graph by calling "apt-cache depends package" for each software package to determine the dependencies between software, and finally establishes a software composition model of the container.
[0082] e) When the dynamic analysis tool finishes the dynamic analysis, it generates an SBOM, which includes information such as the names and versions of the software components used by the container image, as well as the dependencies between the software.
[0083] Dynamic analysis can provide more accurate and comprehensive information about the software components of the container image to identify certain types of vulnerabilities and security issues, which usually only manifest at runtime. Through the SBOM, complete information about the software can be obtained, including component names, license information, version numbers, and suppliers, etc.
[0084] Deduplicate, filter, and clean the data of the software identified by the dynamic analysis, and store the extracted software information and related dependencies in the knowledge graph database. At the same time, after parsing through the SBOM, there are still some software that has not been parsed, then the binary feature matching method is used to parse and extract it into the knowledge graph database, and all the above-parsed software is used as the first layer of the software layer.
[0085] Step 2: Supplement software knowledge and vulnerability knowledge to complete the construction of the final container image knowledge graph: On the basis of parsing the image components, supplement additional software knowledge and vulnerability information in real time by collecting online software repositories (such as maven, PyPI, npm), operating system managers (such as apt, dpkg), and vulnerability libraries (such as NVD, CWE, ULN). To respond to security threats more promptly, use the search engine as the data source and add security events semi-manually at regular intervals to supplement vulnerability information.
[0086] Step 2-1: Parse the online software repository to obtain software knowledge and supplement it to the software layer.
[0087] Based on the analysis of the mirror components, software information and its dependencies in online software repositories (Maven, PyPI, npm), that is, the software library information at the application layer, are collected. At the same time, software information and its dependencies in the operating system manager (such as apt, dpkg), that is, the software library information at the operating system layer, are obtained to solve problems such as missing dependency information and supplement software knowledge in real time. As the subsequent layer of the software layer, it is supplemented to the software layer of the container image knowledge graph.
[0088] Step 2-2: Obtain specific open-source projects, compile and extract to get software knowledge and supplement it to the software layer
[0089] In this embodiment, popular open-source projects are selected from GitHub, and common programming languages such as C / C++, Java, Python, Go, and JavaScript are selected. The top 200 projects of each language in the GitHub open-source project ranking list are collected respectively. After compilation, the binary file features are extracted, and their software information and dependencies are obtained and stored as software in the container image knowledge graph. When there are mirror-related open-source projects, it can avoid the incomplete dependency information of itself and the difficulty in tracking vulnerabilities. It is also used as the subsequent layer of the software layer to supplement the software layer of the container image knowledge graph.
[0090] Step 2-3: Parse the online vulnerability library to obtain vulnerability knowledge, and after matching the vulnerability knowledge with the software parsed from the software bill of materials, form the vulnerability layer of the container image knowledge graph.
[0091] For the software parsed and matched in Step 1-2, there may be CVE vulnerabilities. Usually, one of the main purposes of SBOM is to provide transparency for the components and dependencies of software products, but vulnerability information is not included. Therefore, the vulnerability layer information, including CVSS scores, CWE IDs, and References, etc., is supplemented through an online vulnerability library (such as NVD, CWE, ULN), and the vulnerability information is supplemented and updated regularly. At the same time, it is matched with the software packages and their version numbers parsed from the SBOM to determine the vulnerability attribution.
[0092] Known affected software configurations can be obtained from the vulnerability library, using the CPE standard format, that is, "cpe: <part> : <vendor> : <product> : <version> : <update> : <edition> : <language>”, where the key fields <part>The field can only take three values. 'a' represents the application, 'h' represents the hardware platform, and 'o' represents the operating system. This embodiment only focuses on the information with the value of 'a', that is, the affected scope is the application.
[0093] For example, the CPE field information of the vulnerability "CVE-2022-34750" is "cpe:2.3:a:mediawiki:mediawiki:*:*:*:*:*:*:*:*". Among them, 2.3 is the version of the CPE rule, 'a' indicates that the affected scope is the application, the vendor name is mediawiki, the software name is also mediawiki, and the software version and subsequent information are all '*', which means that all versions are affected by it; usually, getting the software name, software version, and vendor name can be used to match the software packages parsed from the SBOM, and store the vulnerability and the matched software packages in the vulnerability layer of the container image knowledge graph. Since all versions are affected in this example, in this example, the software version information can be ignored for matching.
[0094] Step 2-4: Receive the regularly input security vulnerabilities as vulnerability knowledge and supplement them to the vulnerability layer.
[0095] Due to the lag in the update of security vulnerability data, in order to be able to respond to security threats more timely, relying on the latest information sources on the Internet, security events are added semi-manually regularly to supplement the vulnerability information. The search scope of security events is the security vulnerabilities disclosed on public platforms, such as relevant official accounts, major domestic and foreign network security forum websites, etc. Such information has high timeliness and is not yet included in the online vulnerability database. Considering the frequency of security events, automated collection is performed once a week, and the top ten search results are taken. After manually extracting the information, the security events are roughly classified and the danger level is evaluated. Therefore, the container image knowledge graph regularly receives the input security vulnerabilities, including their rough classification and danger level evaluation, and uses them as supplements to the vulnerability layer.
[0096] It should be noted that the vulnerability knowledge supplemented by the semi-manual method can be further improved and updated when the relevant vulnerabilities in the online vulnerability database are updated.
[0097] Step 3: Based on the image layer, software layer, and vulnerability layer of the container image knowledge graph, quantify the image scope to obtain the set of vulnerabilities corresponding to the images in the to-be-predicted container image repository, and quantify the vulnerability scope to obtain the set of images in the to-be-predicted container image repository affected by specific vulnerabilities.
[0098] Step 3-1: Quantify the image scope: The purpose of this step is to clarify the security profile of a certain image, including the number of existing vulnerabilities, the danger level of vulnerabilities, etc. For example, the image "ubuntu:18.04" contains 8 low-risk vulnerabilities such as "CVE-2022-29458".
[0099] The knowledge graph constructed in this embodiment is generally divided into three layers, from top to bottom are the mirror layer, the software layer, and the vulnerability layer, which is a directed acyclic graph. In the container image knowledge graph, there is a many-to-many relationship between adjacent layers of the mirror and the software, and there is also a dependency relationship between software at the same layer, and there is also a many-to-many relationship between adjacent layers of the software and the vulnerability. Since the system main body is a huge and complex software layer, this embodiment proposes an efficient search algorithm to solve problems such as direct and indirect dependencies of software.
[0100] As Figure 3 shown, 3(a) represents the dependency relationship of a single mirror, and 3(b) represents the actual search path of the algorithm. The search algorithm proposed in this embodiment includes the following steps:
[0101] a) Obtain the specified mirror in the container image repository to be predicted, and find all directly dependent package entities according to the direct dependency relationship between the mirror and the software. This embodiment maintains a dictionary or hash map of package entities that have been visited to avoid repeated searches and make queries and retrievals more efficient.
[0102] b) For each connected package entity in the container image knowledge graph, use the TopologicalSort function of topological sorting to calculate the linear sorting of the package entity according to the dependency relationship of the package entity.
[0103] c) After obtaining the linear sorting, use the Dijkstra algorithm to calculate the shortest path between two package entities.
[0104] d) Apply the BFS algorithm to find all connected vulnerability entities in the shortest path along the shortest path.
[0105] e) Collect all the vulnerability entities found in step d to obtain the final set of vulnerabilities corresponding to the mirror.
[0106] Step 3-2, Vulnerability Scope Quantification: The specific vulnerabilities referred to here can be manually added and updated as described in step 2-4, or a certain high-risk vulnerability can be specified. The purpose of this step is to clarify the scope of influence of a certain vulnerability or security event, including the overall distribution of images affected by the vulnerability, the version range affected by a single image, etc. Among them, the version range affected by a single image, such as the version range affected by the vulnerability "ELSA-2016-3515" for the image mysql includes six versions from "mysql:5" to "mysql:5.7.41-oracle".
[0107] As Figure 4 shown, 4(a) represents the dependency relationship of a single vulnerability, and 4(b) represents the actual search path of the algorithm. Similar to the search algorithm proposed in the image scope quantification, the search algorithm adopted in the vulnerability scope quantification is as follows:
[0108] a) Obtain the specified vulnerability. According to the direct dependency relationship between the vulnerability and the software, find all directly dependent package entities. In this embodiment, a dictionary or hash map of a set of package entities that have been visited is maintained to avoid repeated searches and make querying and retrieval more efficient.
[0109] b) All the package entities accessed by this algorithm are package entities affected by the vulnerability. Use the TopologicalSortReverse function for reverse topological sorting of the package entities, and calculate the reverse linear sorting of the package entities according to the dependency relationship of the package entities.
[0110] c) After obtaining the reverse linear sorting, use the reverse Dijkstra algorithm to calculate the shortest path between two package entities.
[0111] d) Apply the reverse BFS algorithm to find all connected mirror entities in the shortest path along the shortest path.
[0112] e) Collect all the mirror entities found in step d to obtain the final set of mirrors affected by the vulnerability.
[0113] Step 4. Output the risk prediction result of the container image repository to be predicted according to the vulnerability set, the mirror set, and the mirror download volume.
[0114] In this embodiment, entity types are defined according to the objects involved in the container image knowledge graph, including images, software, vulnerabilities, etc. A triple T=(I, P, V) is defined, where I is the node type of the image (Image), P is the node type of the software (Package), and V is the node type of the vulnerability (Vulnerability).
[0115] Define I P as the mirror download volume. The download volume range of the popular mirrors selected in this embodiment ranges from one hundred thousand to ten billion magnitudes; P V is the proportion of affected software in the mirror, that is, the ratio of the number of affected software in the mirror to the total number of software; S V is the threat score of the vulnerability. Here, the CVSS score is adopted. Most vulnerability scores adopt the CVSS 3.x version, preferably version 3.1, and secondarily version 3.0. For some old vulnerabilities, such as CVE-2007-0994, if there is only a score of version 2.0, then version 2.0 is selected.
[0116] Step 4-1. Security assessment of mirror download volume: Since the range of the mirror download volume I P has a too large gap, perform normalization processing on I P using the log function conversion. Based on Formula 1, obtain the security score S I of the mirror download volume, and the range is between (0, 1).
[0117]
[0118] Wherein, I P is the mirror download volume, and I Pmax is the maximum download volume of the images in the container image repository to be predicted.
[0119] Step 4-2, Software security assessment.
[0120] First, obtain the total number of software in the image based on the vulnerability set, that is, through the mirror scope quantification process in Step 3-1, count the total number of software PSUM in the image, which is all the directly dependent package entities found.
[0121] Then, obtain the number of software affected by vulnerabilities in the image based on the image set, that is, through the vulnerability scope quantification process in Step 3-2, count the number of affected software PC in the image, which is the package entity that has a direct dependency on any specified vulnerability and belongs to the currently evaluated image.
[0122] Therefore, the proportion P V of the affected software in the image is the ratio of PC to PSUM, and it is normalized using the maximum-minimum normalization. Based on Formula 2, the software security score S P is obtained, and the range is between (0, 1).
[0123]
[0124] Wherein, P V is the proportion of the affected software in the image, P Vmin is the minimum proportion of the affected software in the images in the container image repository to be predicted, and P Vmax is the maximum proportion of the affected software in the images in the container image repository to be predicted.
[0125] Step 4-3, Vulnerability threat score assessment: The vulnerability set obtained from the mirror scope quantification process in Step 3-1 contains the score situation of the vulnerabilities. Based on Formula 3, the weighted sum vulnerability score V CS of the corresponding vulnerability set of the image is obtained, where V CH is the score of high-risk vulnerabilities (for example, the CVSS score is higher than 7.0, which can be adjusted), and V CL is the score of low-risk vulnerabilities (for example, the CVSS score is lower than 7.0, which can be adjusted). Weights a and b are assigned to the two respectively. In this embodiment, a = 0.7 and b = 0.3 are adopted, and in other embodiments, they can be adjusted according to actual needs. The maximum-minimum normalization is used to normalize VCS, and based on Formula 4, the vulnerability threat score S V is obtained, and the range is between (0, 1).
[0126]
[0127]
[0128] In the formula, V CS is the weighted sum vulnerability score of the vulnerability set corresponding to the mirror, n h is the number of high-risk vulnerabilities in the vulnerability set corresponding to the mirror, n l is the number of low-risk vulnerabilities in the vulnerability set corresponding to the mirror, is the score of the i-th high-risk vulnerability, is the score of the j-th low-risk vulnerability, a is the weight of high-risk vulnerabilities, b is the weight of low-risk vulnerabilities, S V is the vulnerability threat score, is the minimum weighted sum vulnerability score in the container image repository to be predicted, is the maximum weighted sum vulnerability score in the container image repository to be predicted.
[0129] Step 4-4, Image Security Assessment: For S I , S P and S V obtained from Step 4-1, Step 4-2, and Step 4-3, weights α, β, and γ are respectively assigned, and the image score S T is obtained based on Formula 5, and normalized using maximum-minimum normalization. The final image security score S SUM is obtained based on Formula 6, with a range between (0, 10). Generally, the higher the image download volume I P , the smaller the proportion P V of affected software, the lower the vulnerability threat score S V , and the higher the image security score S SUM . The higher the image security score S SUM , the higher the image security, and vice versa, the lower the security.
[0130] S T = α * S I - β * S P - γ * S V (5)
[0131]
[0132] In the formula, α is the weight of S I , β is the weight of S P , γ is the weight of S V , S Tmin is the minimum image score in the container image repository to be predicted, S Tmax is the maximum image score in the container image repository to be predicted.
[0133] Step 4-5, Security Evaluation of Image Repository: The risk prediction result of the container image repository to be predicted is obtained based on the image security scores of each image in the container image repository.
[0134] It should be noted that obtaining the risk prediction result based on the image security score can be to output the risk prediction result according to predefined rules, or to input the image security score into a pre-trained risk prediction model (deep learning model) to obtain the risk prediction result, which is not limited in this embodiment. The obtained risk prediction result can be used for container image repository optimization, screening, etc.
[0135] For the sake of easy understanding, a process of outputting the risk prediction result according to predefined rules in this embodiment is as follows: Define that the image security score of 0-3 is a high-risk image, 4-7 is a medium-risk image, and 8-10 is a safe image. Relevant image sets can be selected according to the scores and additional marks can be added to them. For the container image repository as a whole, it can be known the number and proportion of images in different scoring intervals. The larger the proportion of high-risk images, the lower the overall security of the container image repository, and vice versa, the larger the proportion of safe images, the higher the overall security of the container image repository.
[0136] At the same time, screening can be additionally performed according to the CVSS score of the vulnerability. For example, vulnerabilities with a CVSS score higher than 8.0 are selected as high-risk vulnerabilities, and low-risk vulnerabilities are omitted. The affected images are screened out through the vulnerability scope quantification process in Step 3-2, and their security scores are iteratively calculated to screen out the high-risk image set. Similarly, calculations can be performed for vulnerabilities with a CVSS score higher than 6.0, 7.0, etc., and the distribution trend of high-risk images affected by vulnerabilities of different risk levels can be observed. Similarly, vulnerabilities with a CVSS score within a certain range can be screened, such as scores of 6.0-7.0, 7.0-8.0, etc., and the concentrated distribution of high-risk images affected by vulnerabilities within a certain CVSS score range can be observed.
[0137] In another embodiment, the present invention also provides a risk prediction device for a container image repository based on a knowledge graph, including a processor and a memory storing a number of computer instructions. When the computer instructions are executed by the processor, the steps of the risk prediction method for the container image repository based on the knowledge graph are implemented.
[0138] For the specific limitations of the risk prediction device for the container image repository based on the knowledge graph, reference can be made to the limitations of the risk prediction method for the container image repository based on the knowledge graph described above, which will not be elaborated here.
[0139] The memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can run on the processor, and the processor realizes the method for predicting risks of a container image repository based on a knowledge graph in the embodiments of the present invention by running the computer program stored in the memory.
[0140] Among them, the memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0141] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0142] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0143] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.< / part> < / language> < / edition> < / update> < / version> < / product> < / vendor> < / part>
Claims
1. A risk prediction method for container image repositories based on a knowledge graph, characterized in that The method for predicting risks of a container image repository based on a knowledge graph includes: Step 1: Construct a basic container image knowledge graph; Step 1-1: Obtain original image data, extract image entities and their attributes and relationships from the original image data to obtain valid images, and form the image layer of the container image knowledge graph; Step 1-2: Use a dynamic analysis method to parse the valid images, construct a software bill of materials, and use the parsed software as the software layer of the container image knowledge graph; Step 2: Supplement software knowledge and vulnerability knowledge to complete the construction of the final container image knowledge graph; Step 2-1: Parse the online software repository to obtain software knowledge and supplement it to the software layer; Step 2-2: Obtain specific open-source projects, compile and extract to obtain software knowledge and supplement it to the software layer; Step 2-3: Parse the online vulnerability database to obtain vulnerability knowledge. After matching the vulnerability knowledge with the software parsed in the software bill of materials, form the vulnerability layer of the container image knowledge graph; Step 2-4: Receive regularly input security vulnerabilities as vulnerability knowledge and supplement them to the vulnerability layer; Step 3: Based on the image layer, software layer and vulnerability layer of the container image knowledge graph, quantify the image scope to obtain the set of vulnerabilities corresponding to the images in the container image repository to be predicted, and quantify the vulnerability scope to obtain the set of images affected by specific vulnerabilities in the container image repository to be predicted; Step 4: Output the risk prediction result of the container image repository to be predicted according to the set of vulnerabilities, the set of images and the image download volume.
2. The method for predicting risks of a container image repository based on a knowledge graph according to claim 1, wherein The obtaining of the valid images includes: Perform knowledge fusion and disambiguation on the extracted image entities and their attributes and relationships, and discard the images that do not contain the SHA256 attribute to obtain valid images.
3. The method for predicting risks in a container image repository based on a knowledge graph according to claim 1, wherein The using of the dynamic analysis method to parse the valid images and construct a software bill of materials includes: Load the container running environment; Run the container image in the container running environment, and start a dynamic analysis tool to monitor the behavior of the container image, obtain a list of all installed software components of the container image currently, and identify the software components loaded and executed by the container image; Determine the dependency relationships existing between software for each software component, and establish a software composition model of the container; When the dynamic analysis tool finishes dynamic analysis, output the software bill of materials, and the software bill of materials includes the names, versions of the software components used by the container image, and the dependency relationships between the software.
4. The method for predicting risks of a container image repository based on a knowledge graph according to claim 1, characterized in that, The quantifying of the image scope to obtain the set of vulnerabilities corresponding to the images in the container image repository to be predicted includes: Obtain images, and find all directly dependent package entities according to the direct dependency relationship between the images and the software; For each package entity, calculate the linear sorting of the package entity according to the dependency relationship of the package entity in the container image knowledge graph; After obtaining the linear sorting, calculate the shortest path between two package entities; Find all connected vulnerability entities in the shortest path along the shortest path; Collect all the discovered vulnerability entities to obtain the final set of vulnerabilities corresponding to the images.
5. The method for predicting risks of a container image repository based on a knowledge graph according to claim 1, wherein, The quantifying of the vulnerability scope to obtain the set of images affected by specific vulnerabilities in the container image repository to be predicted includes: Obtain vulnerabilities, and find all directly dependent package entities according to the direct dependency relationship between the vulnerabilities and the software; For each package entity, calculate the reverse linear ordering of the package entity according to the dependency relationship of the package entity in the container image knowledge graph; After obtaining the reverse linear ordering, calculate the shortest path between two package entities; Find all connected image entities in the shortest path along the shortest path; Collect all the discovered image entities to obtain the final set of images affected by the vulnerability.
6. The method for predicting risks of a container image repository based on a knowledge graph according to claim 1, wherein Output the risk prediction result of the container image repository to be predicted according to the vulnerability set, the image set, and the image download volume, including: The risk prediction result of the container image repository to be predicted is obtained according to the image security scores of each image in the container image repository, and the image security scores are obtained according to the vulnerability set, the image set, and the image download volume.
7. The method for predicting risks of a container image repository based on a knowledge graph according to claim 6, wherein The calculation formula of the image security score is as follows: S T = α * S I - β * S P - γ * S V Where S T is the mirror score, S I is the security score of the mirror download volume, α is the weight of S I , S P is the software security score, β is the weight of S P , S V is the vulnerability threat score, γ is the weight of S V , S SUM is the mirror security score, S Tmin is the minimum mirror score in the container image repository to be predicted, S Tmax is the maximum mirror score in the container image repository to be predicted.
8. The method for predicting risks of a container image repository based on a knowledge graph according to claim 7, wherein, The calculation formula of the image download volume security score is as follows: Where, I P is the mirror download volume, and I Pmax is the maximum download volume of the mirrors in the container image repository to be predicted.
9. The method for predicting risks of a container image repository based on a knowledge graph according to claim 7, wherein The calculation process of the software security score is as follows: Obtain the total number of software in the image based on the vulnerability set; Obtain the number of software affected by vulnerabilities in the image based on the image set; Obtain the proportion of affected software according to the number of software affected by vulnerabilities and the total number of software; Calculate the software security score as follows: Wherein, P V is the proportion of affected software in the mirror, and P Vmin is the minimum proportion of affected software in the images in the container image repository to be predicted, and P Vmax is the maximum proportion of affected software in the images in the container image repository to be predicted.
10. The method for predicting risks of a container image repository based on a knowledge graph according to claim 7, wherein The calculation formula of the vulnerability threat score is as follows: Where V CS is the weighted sum vulnerability score of the vulnerability set corresponding to the mirror, n h is the number of high-risk vulnerabilities in the vulnerability set corresponding to the mirror, n l is the number of low-risk vulnerabilities in the vulnerability set corresponding to the mirror, is the score of the i-th high-risk vulnerability, is the score of the j-th low-risk vulnerability, a is the weight of high-risk vulnerabilities, b is the weight of low-risk vulnerabilities, S V is the vulnerability threat score, is the minimum weighted sum vulnerability score in the container image repository to be predicted, is the maximum weighted sum vulnerability score in the container image repository to be predicted.
Citation Information
Patent Citations
Template representation-based container mirror image configuration file automatic generation method and device
CN116578389A