Open source vulnerability analysis method and device and computer readable storage medium

By generating a security vulnerability knowledge graph and using a graph search algorithm, the problem of not being able to identify the vulnerability propagation path between the dependency levels of open source software in the existing technology is solved, and efficient and accurate vulnerability analysis and troubleshooting is achieved.

CN120012090APending Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311523135.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing open source vulnerability analysis system cannot identify the propagation path between the dependency levels of open source software, resulting in inefficient vulnerability investigation and high cost of troubleshooting.

Method used

By generating a security vulnerability knowledge graph, including vulnerability nodes, software nodes, call relationships and dependencies, the graph search algorithm is used to scan the accessible paths from vulnerability nodes to software nodes to generate vulnerability analysis results.

Benefits of technology

Accurately identify the propagation paths of open source vulnerabilities between dependency levels, reduce false positives, improve the accuracy of vulnerability discovery, and reduce resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012090A_ABST
    Figure CN120012090A_ABST
Patent Text Reader

Abstract

The invention discloses an open source vulnerability analysis method. The method comprises the steps of generating a security vulnerability knowledge graph of a to-be-analyzed software product; the security vulnerability knowledge graph comprises vulnerability nodes, software nodes of the to-be-analyzed software product, a calling relationship between the to-be-analyzed software product and open source software, and a dependency relationship between the open source software; the calling relationship comprises a calling relationship between functions in the software code; scanning a reachable path from the vulnerability node to the software node in the security vulnerability knowledge graph based on a graph search algorithm; and generating a vulnerability analysis result based on the scanning condition. Open source vulnerability analysis is carried out through the security vulnerability knowledge graph, and the technical problems of low vulnerability investigation efficiency and high investigation cost caused by the phenomenon that propagation paths between dependency levels of open source software cannot be identified and vulnerability analysis has false alarms can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an open source vulnerability analysis method, an open source vulnerability analysis device, and a computer-readable storage medium. Background Art

[0002] Open source software is an indispensable part of the modern information industry, and the security risks brought by open source software are becoming increasingly severe. Open source software vulnerability scanning is an automated security tool used to detect and identify possible security vulnerabilities and vulnerability risks in open source software. At present, the detection of open source vulnerabilities in software products is mainly based on the method of matching software names and version numbers.

[0003] Open source vulnerability analysis systems generally issue vulnerability alerts by comparing the open source software name and version number strings with the open source software versions affected by the vulnerabilities. However, there are problems such as open source vulnerability analysis systems not considering the risk of vulnerability propagation in the propagation chain, being unable to identify the propagation paths between open source software dependency levels, and false positives in vulnerability analysis, resulting in low vulnerability detection efficiency and high detection costs. Summary of the invention

[0004] The present application provides an open source vulnerability analysis method, device and computer-readable storage medium, which performs open source vulnerability analysis through a security vulnerability knowledge graph, and can solve the technical problems of being unable to identify the propagation paths between dependency levels of open source software and the existence of false positives in vulnerability analysis, resulting in low vulnerability detection efficiency and high detection costs.

[0005] In a first aspect, the present application provides an open source vulnerability analysis method, which is applied to an open source vulnerability analysis device or an open source vulnerability analysis system, and the device or system may include a computing device such as a server; the method includes:

[0006] Generate a security vulnerability knowledge graph of the software product to be analyzed; the security vulnerability knowledge graph includes vulnerability nodes, software nodes of the software product to be analyzed, calling relationships between the software product to be analyzed and open source software, and dependency relationships between open source software; the calling relationships and dependency relationships include calling relationships between functions in the software code;

[0007] Scan the accessible path from the vulnerability node to the above software node in the security vulnerability knowledge graph based on the graph search algorithm;

[0008] Generate vulnerability analysis results based on the scan results.

[0009] The security vulnerability knowledge graph generated by the above embodiment covers all calling relationships of software products, including the dependency relationships between software products and open source software, within open source software, and between open source software. It can fully reflect the propagation path of open source software vulnerabilities between open source software dependency levels, thereby accurately determining whether open source vulnerabilities will affect software products, and pushing vulnerability information that truly affects application software projects to developers for rectification, reducing the interference caused by invalid components and vulnerability information, and reducing resource consumption caused by false alarms while improving the accuracy of security vulnerability discovery.

[0010] In a possible implementation, the security vulnerability knowledge graph may be updated at a preset period.

[0011] By regularly updating the security vulnerability knowledge graph, the security vulnerability analysis of software products can be completed more comprehensively and accurately. In particular, for released software products that have undergone security vulnerability analysis in the early stage, security vulnerability analysis can be regularly performed based on the updated security vulnerability knowledge graph in the later stage, and leakage analysis and warning can be continuously performed to ensure the normal operation of software products.

[0012] In a possible implementation, the above-mentioned generation of the security vulnerability knowledge graph of the software product to be analyzed includes:

[0013] Obtaining a list of open source software dependencies for the software product to be analyzed; the list of open source software dependencies includes multiple open source software that the software product to be analyzed depends on;

[0014] Generate a function call graph for each open source software in the above open source software dependency list;

[0015] The generated function call graphs are merged to generate a software dependency entity; the software dependency entity includes the dependency relationship between the open source software, and the dependency relationship between the open source software includes the dependency relationship of each function call graph.

[0016] Through the above embodiments, based on the open source software dependency list obtained from different vulnerability databases, open source software with direct dependencies and open source software with indirect dependencies can be obtained; constructing the calling relationship between software products and open source software at the function granularity can accurately locate the calling relationship between source codes, ensure the accuracy of security vulnerability discovery, and improve the traditional use of open source software name and version number matching to determine whether open source vulnerabilities affect software products. There are false alarms and waste of resources.

[0017] In a possible implementation, the generating of the function call graph of each open source software in the open source software dependency list includes:

[0018] For each open source software in the above open source software dependency list, decompose each open source code file in the open source software;

[0019] For each open source code file in the above open source software, decompose each class in the open source code file;

[0020] For each class code snippet in the above open source code file, decompose each function in the class;

[0021] For each of the above function code snippets, decompose the statements contained in the function;

[0022] For each statement in the same function, a control flow graph is generated with the statement as a node; if the statement has the behavior of calling other functions, the above control flow graph contains the call node saved based on the structured cache, and there is an edge between the statement node that calls other functions and the above call node;

[0023] Generate a function call graph within the same class based on the above control flow graph;

[0024] Merge the function call graphs of all classes in the same open source code file to generate a function call graph of the open source code file;

[0025] The function call graphs of each open source code file in the same open source software are merged to generate a respective function call graph for each open source software in the above open source software dependency list.

[0026] Through the above embodiments, the dependencies between software products and open source software, within open source software, and between open source software are taken into consideration, and the function call graph, forward control flow graph, and backward control flow graph in the source code are mined using intra-process analysis, inter-process analysis, and dependency analysis. The semantic and grammatical information of the source code is retained at the function granularity, and the use of function call graphs and control flow graphs to accurately characterize the dependencies of open source software is able to not only accurately locate the calling relationships between source codes, but also narrow the scope of dependencies between software products and open source software, and identify the open source software actually used by the application software project.

[0027] In a possible implementation, for each statement in the same function, a control flow graph is generated with the statement as a node, including:

[0028] Divide the function into statements, with each statement as a statement node;

[0029] If the statement contains the behavior of calling other functions, the calling node is determined, and the calling node is saved based on the structured cache;

[0030] A control flow graph is generated based on statement nodes and call nodes; the edges between nodes in the control flow graph represent the execution order of the nodes or the call relationship between the nodes.

[0031] Through the above-mentioned embodiments, based on the analysis within the open source software process, a control flow graph is generated according to the execution order and calling relationship within the software code, and the semantic grammatical information of the source code is retained at the function granularity, which can ensure that the function call graph and the control flow graph are accurate in representing the open source software dependency relationship.

[0032] In a possible implementation, the above-mentioned determination of the calling node includes: if there is a hidden node, traversing upward with the hidden node as a leaf node, and taking the found root node as the calling node; and / or,

[0033] Different data dependency analyses are performed on calls with different parameters to determine the call nodes.

[0034] Through the above embodiment, the relationship between hidden nodes is taken into consideration. By traversing upward with the hidden nodes as leaf nodes to find the root node, the real calling node can be found. In addition, data dependency analysis is performed on the existing calling nodes. For example, based on the difference in the calling context, different analyses are performed on calls with different parameters. By analyzing the same function multiple times, the accuracy of the vulnerability analysis can be further guaranteed.

[0035] In a possible implementation, the above-mentioned generating a function call graph within the same class based on the above-mentioned control flow graph; merging the function call graphs of all classes in the same open source code file to generate a function call graph of the open source code file; including:

[0036] Get all function blocks in the same open source code file;

[0037] Based on the control flow graph, an edge is generated between the first root node and the first leaf node (from the first root node to the first leaf node), an edge is generated between the second leaf node and the third leaf node (from the second leaf node to the third leaf node), and an edge is generated between the fourth leaf node and the second root node (from the fourth leaf node to the second root node);

[0038] Among them, the function block with a class node in the upper layer is the above-mentioned first leaf node, and the class node with a function block below is the above-mentioned first root node; the function block that calls other functions is the above-mentioned second leaf node, and the function block called by other functions is the above-mentioned third leaf node; the function block that has no class node in the upper layer and has a calling relationship with the function in the class is the fourth leaf node, and the class node called by the above-mentioned fourth leaf node is the above-mentioned second root node;

[0039] If the called function is not in the open source code file, the called function is saved as a calling node based on the structured cache.

[0040] In a possible implementation, the function call graphs of each open source code file in the same open source software are merged to generate a function call graph for each open source software in the open source software dependency list, including:

[0041] For the above-mentioned calling node, determining the target source code file containing the above-mentioned calling node from other open source code files in the above-mentioned same open source software;

[0042] Generate an import node for the above target source code file;

[0043] An abstract syntax tree is generated based on the import node. If there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted.

[0044] In a possible implementation, the above-mentioned function call graphs are merged to generate a software dependency entity, including:

[0045] Based on the function call graph generated by each open source software, for the open source software with call nodes, extract the import nodes from other open source software in the dependency list of the open source software;

[0046] An abstract syntax tree is generated based on the extracted import node. If there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted; until all the call nodes in the open source software with the call node are processed.

[0047] Through the above-mentioned embodiments, software dependency entities can be generated efficiently and completely, taking into account the dependency relationships between software products and open source software, within open source software, and between open source software. Intra-process analysis, inter-process analysis, and dependency analysis are used to mine function call graphs, forward control flow graphs, and backward control flow graphs in the source code. The calling relationships between source codes can be accurately located, and the dependency scope between software products and open source software can be narrowed, and the open source software actually used by the application software project can be identified.

[0048] In a possible implementation, the generating of the security vulnerability knowledge graph of the software product to be analyzed further includes:

[0049] Associating a vulnerability node where a software vulnerability exists with the above software dependency entity;

[0050] Based on the execution order involved in the above vulnerability nodes, the calling relationship of each node in the above security vulnerability knowledge graph is integrated.

[0051] Through the above-mentioned embodiments, vulnerability nodes with software vulnerabilities are associated with software dependency entities, and the execution order of nodes involved in the vulnerability nodes is considered, including the execution order of vulnerability file nodes, vulnerability class nodes, vulnerability function nodes and vulnerability statement nodes, etc., so as to comprehensively characterize the calling relationship of each node in the security vulnerability knowledge graph and ensure the accuracy of vulnerability analysis.

[0052] In a possible implementation, before associating the vulnerability node with the software vulnerability on the software dependency entity, the method further includes:

[0053] For each vulnerable open source software, obtain the open source vulnerability source code fragment of the vulnerable open source software;

[0054] Analyze the open source vulnerability source code fragments to generate a vulnerability knowledge graph of the open source software; the vulnerability knowledge graph includes the name information, version information, vulnerability statements and vulnerability propagation path of the open source software, and the vulnerability propagation path includes the propagation path from the root node to the node of the vulnerability statement;

[0055] Integrate the vulnerability knowledge graphs of multiple (or all) vulnerable open source software.

[0056] Through the above embodiments, multi-layer hierarchical vulnerability source code root location is achieved, and the propagation of open source vulnerabilities can be narrowed down from the entire open source software to a certain class or function method in the open source software; in the process of integrating the vulnerability knowledge graphs of the respective vulnerable open source software, it is also possible to perform unified analysis on heterogeneous data with different structures from different vulnerability libraries, and uniformly normalize and generate a vulnerability knowledge graph in a fixed format, thereby solving the problem of isolated and scattered vulnerability information and reducing the workload of manual labeling.

[0057] In a possible implementation, obtaining the open source vulnerability source code fragment of the vulnerable open source software includes:

[0058] Based on the open source vulnerability library, vulnerability knowledge is extracted from the unstructured text in the open source vulnerability web pages;

[0059] Extract open source vulnerability source code fragments of vulnerable open source software from the vulnerability knowledge.

[0060] Through the above embodiments, it is possible to automatically extract structured data from unstructured text in the industry vulnerability library and deeply mine vulnerability source code data in the unstructured text.

[0061] In a possible implementation, analyzing the open source vulnerability source code fragment includes:

[0062] Slicing the open source vulnerability source code fragment, parsing the source code fragment into a syntax node of an abstract syntax tree, adding node labels to the nodes of the abstract syntax tree, mapping the source code fragment to the abstract syntax tree, and adding instance nodes to the abstract syntax tree;

[0063] The path of the leaf nodes of the abstract syntax tree is traversed to generate a propagation path from the root node to the leaf nodes; the leaf nodes are label nodes marked with open source vulnerabilities.

[0064] Through the above embodiments, the propagation path of the open source vulnerability within the open source software can be accurately generated, ensuring that whether the vulnerability affects the application software project can be accurately located on the call relationship chain.

[0065] In a possible implementation, after extracting vulnerability knowledge from unstructured text in an open source vulnerability web page based on an open source vulnerability library, the method further includes:

[0066] Extracting an open source vulnerability description statement of the vulnerable open source software from the vulnerability knowledge;

[0067] Performing text parsing on the open source vulnerability description statement to obtain vulnerability named entity information; the vulnerability named entity information includes the name information, version information, vulnerability statement of the vulnerable open source software and at least one of the following: vulnerability file information, vulnerability class, and vulnerability function;

[0068] Wherein, in the process of generating the vulnerability knowledge graph of the vulnerable open source software, the source code information extracted by analyzing the open source vulnerability source code fragment is integrated with the vulnerability named entity information.

[0069] Through the above embodiment, the vulnerability named entity information of the open source vulnerability software is obtained through the preset vulnerability named entity recognition model, which can ensure that complete vulnerability information is obtained and the accuracy of vulnerability analysis is guaranteed.

[0070] In a possible implementation, the text parsing of the open source vulnerability description statement to obtain vulnerability named entity information includes:

[0071] Performing word segmentation processing on the open source vulnerability description sentence to obtain a sequence after word segmentation;

[0072] The sequence after word segmentation is processed through a bidirectional encoder representation network based on Transformer to obtain word encoding, sentence encoding and position encoding;

[0073] The word code, the sentence code and the position code are summed and combined;

[0074] The merged vector encoding is processed through a bidirectional long short-term memory model and conditional random field to output vulnerability named entity information.

[0075] Through the above embodiments, vulnerability named entity information of open source vulnerability software can be obtained quickly and efficiently.

[0076] In summary, the embodiment of the present application converts the vulnerability propagation path into a directed graph analysis, and uses graph theory knowledge to automatically analyze the propagation path of open source vulnerabilities between open source software dependency levels by constructing a security vulnerability knowledge graph. The security vulnerability knowledge graph covers all calling relationships between software products and open source software, and can accurately locate whether the vulnerability affects the application software project on the calling relationship chain, and push the vulnerability information that really affects the application software project to the developers for rectification, thereby reducing the interference caused by invalid components and vulnerability information, and reducing the resource consumption caused by false alarms while improving the accuracy of security vulnerability discovery; in addition, the embodiment of the present application is language-independent and is applicable to vulnerability information in any format and all high-level languages; it can realize a fully automated extraction and analysis process, and through a series of operations such as open source dependency list extraction, vulnerability knowledge extraction, open source dependency graph generation, and security vulnerability knowledge graph mining of vulnerability propagation paths, the efficiency of open source vulnerability warnings and risk assessments in software products is improved.

[0077] The embodiments of the present application can be applied to the current technical solution of performing the open source vulnerability analysis after the software product is compiled and generated and before it is released; it can also directly perform the open source vulnerability analysis on the uncompiled software product at the premise of sacrificing some performance, or it can be the open source software dependency demand analysis stage, or it can be deployed right after the release, so as to realize continuous vulnerability scanning of the software product and real-time feedback of vulnerability warning information.

[0078] In a second aspect, the present application provides an open source vulnerability analysis device, including:

[0079] A graph generation unit, used to generate a security vulnerability knowledge graph of the software product to be analyzed; the security vulnerability knowledge graph includes vulnerability nodes, software nodes of the software product to be analyzed, calling relationships between the software product to be analyzed and open source software, and dependency relationships between open source software; the calling relationships include calling relationships between functions in the software code;

[0080] A path scanning unit, used for scanning a reachable path from a vulnerability node to the software node in a security vulnerability knowledge graph based on a graph search algorithm;

[0081] The result generating unit is used to generate vulnerability analysis results based on the scanning situation.

[0082] In a possible implementation, the graph generation unit includes:

[0083] An acquisition unit, configured to acquire an open source software dependency list for the software product to be analyzed; the open source software dependency list includes a plurality of open source software that the software product to be analyzed depends on;

[0084] A first generating unit, configured to generate a function call graph for each open source software in the open source software dependency list;

[0085] The second generating unit is used to merge the generated function call graphs to generate a software dependency entity; the software dependency entity includes the dependency relationship between the open source software, and the dependency relationship between the open source software includes the dependency relationship of each function call graph.

[0086] In a possible implementation manner, the first generating unit includes:

[0087] The generating of a function call graph for each open source software in the open source software dependency list includes:

[0088] A first decomposition unit is used to decompose each open source code file in the open source software for each open source software in the open source software dependency list;

[0089] A second decomposition unit is used to decompose each open source code file in the open source software into each class in the open source code file;

[0090] A third decomposition unit is used to decompose each function in the class according to the code snippet of each class in the open source code file;

[0091] A fourth decomposition unit is used to decompose the code snippet of each function into statements contained in the function;

[0092] A control flow graph generation unit is used to generate a control flow graph for each statement in the same function with the statement as a node; if the statement has the behavior of calling other functions, the control flow graph includes a call node saved based on the structured cache, and there is an edge between the statement node that calls other functions and the call node;

[0093] A first call graph generating unit, configured to generate a function call graph within the same class based on the control flow graph;

[0094] A second call graph generating unit is used to merge the function call graphs of all classes in the same open source code file to generate a function call graph of the open source code file;

[0095] The first merging unit is used to merge the function call graphs of each open source code file in the same open source software to generate a respective function call graph for each open source software in the open source software dependency list.

[0096] In a possible implementation, the control flow graph generating unit includes:

[0097] A partitioning unit is used to partition a function into statements, with each statement being a statement node;

[0098] A first determining unit is used to determine a calling node if the statement has a behavior of calling other functions, and save the calling node based on the structured cache;

[0099] A generation subunit is used to generate a control flow graph based on statement nodes and call nodes; the edges between nodes in the control flow graph represent the execution order of the nodes or the call relationship between the nodes.

[0100] In a possible implementation manner, the determining unit includes:

[0101] A traversal unit, used to traverse upwards the hidden node as a leaf node if there is a hidden node, and use the found root node as a calling node; and / or,

[0102] The data dependency analysis unit is used to perform different data dependency analyses on calls with different parameters to determine the call node.

[0103] In a possible implementation, the second call graph generating unit includes:

[0104] A function acquisition unit, used to acquire all function blocks in the same open source code file;

[0105] An edge generation unit, configured to generate, based on the control flow graph, an edge between the first root node and the first leaf node, an edge between the second leaf node and the third leaf node, and an edge between the fourth leaf node and the second root node;

[0106] Among them, the function block with a class node in the upper layer is the first leaf node, and the class node with a function block below is the first root node; the function block that calls other functions is the second leaf node, and the function block called by other functions is the third leaf node; the function block that has no class node in the upper layer and has a calling relationship with the function in the class is the fourth leaf node, and the class node called by the fourth leaf node is the second root node;

[0107] The saving unit is used to save the called function as a calling node based on the structured cache if the called function is not in the open source code file.

[0108] In a possible implementation, the first merging unit includes:

[0109] A second determining unit is used to determine, for the calling node, a target source code file containing the calling node from other open source code files in the same open source software;

[0110] An import node generation unit, used to generate an import node of the target source code file;

[0111] The first abstract syntax tree generation unit is used to generate an abstract syntax tree based on the import node, if there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted.

[0112] In a possible implementation manner, the second generating unit includes:

[0113] An import node extraction unit is used to extract import nodes from other open source software in the dependency list of the open source software for the open source software having call nodes based on the function call graph generated by each open source software;

[0114] A second abstract syntax tree generation unit is used to generate an abstract syntax tree based on the extracted import node. If there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted; until all the call nodes in the open source software with the call node are processed.

[0115] In a possible implementation, the graph generation unit further includes:

[0116] An associating unit, used to associate a vulnerability node having a software vulnerability with the software dependency entity;

[0117] The first integration unit is used to integrate the calling relationship of each node in the security vulnerability knowledge graph based on the execution order involved in the vulnerability node.

[0118] In a possible implementation, the device further includes:

[0119] A source code fragment acquisition unit is used to acquire, for each vulnerable open source software, an open source vulnerability source code fragment of the vulnerable open source software before associating the vulnerability node with the software vulnerability on the software dependency entity;

[0120] A source code fragment analysis unit, configured to analyze the open source vulnerability source code fragment and generate a vulnerability knowledge graph of the vulnerable open source software; the vulnerability knowledge graph includes name information, version information, vulnerability statements, and vulnerability propagation paths of the vulnerable open source software, wherein the vulnerability propagation path includes a propagation path from a root node to a node of the vulnerability statement;

[0121] The second integration unit is used to integrate the vulnerability knowledge graphs of multiple (or all) vulnerable open source software.

[0122] In a possible implementation, the source code segment obtaining unit includes:

[0123] A vulnerability knowledge extraction unit is used to extract vulnerability knowledge from unstructured text in an open source vulnerability web page based on an open source vulnerability library;

[0124] The extraction subunit is used to extract the open source vulnerability source code fragment of the vulnerable open source software from the vulnerability knowledge.

[0125] In a possible implementation, the source code segment analysis unit includes:

[0126] A slicing unit, configured to slice the open source vulnerability source code fragment, parse the source code fragment into a syntax node of an abstract syntax tree, add node labels to the nodes of the abstract syntax tree, map the source code fragment to the abstract syntax tree, and add instance nodes to the abstract syntax tree;

[0127] A path generation unit is used to traverse the path of the leaf nodes of the abstract syntax tree and generate a propagation path from the root node to the leaf node; the leaf node is a label node marked with an open source vulnerability.

[0128] In a possible implementation, the device further includes:

[0129] A description sentence extraction unit, used to extract an open source vulnerability description sentence of the vulnerable open source software from the vulnerability knowledge;

[0130] A text parsing unit, configured to perform text parsing on the open source vulnerability description statement to obtain vulnerability named entity information; the vulnerability named entity information includes the name information, version information, vulnerability statement of the vulnerable open source software, and at least one of the following: vulnerability file information, vulnerability class, and vulnerability function;

[0131] Wherein, in the process of generating the vulnerability knowledge graph of the vulnerable open source software, the source code information extracted by analyzing the open source vulnerability source code fragment is integrated with the vulnerability named entity information.

[0132] In a possible implementation, the text parsing unit includes:

[0133] A word segmentation processing unit, used to perform word segmentation processing on the open source vulnerability description sentence to obtain a sequence after word segmentation;

[0134] The network processing unit is used to process the sequence after word segmentation through a bidirectional encoder representation network based on Transformer to obtain word encoding, sentence encoding and position encoding;

[0135] A summing and merging unit, used for summing and merging the word code, the sentence code and the position code;

[0136] The entity information output unit is used to process the merged vector encoding through a bidirectional long short-term memory model and a conditional random field, and output vulnerability named entity information.

[0137] In a third aspect, the present application provides an open source vulnerability analysis device, comprising: one or more processors and one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer executable programs. When the one or more processors execute the computer executable programs, the electronic device executes any possible implementation method in the first aspect.

[0138] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor of an electronic device, the electronic device executes any possible implementation method in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0139] Figure 1 This is an application scenario diagram of an open source vulnerability analysis system provided by this application;

[0140] Figure 2 It is an architecture diagram of an open source vulnerability analysis system provided by this application;

[0141] Figure 3 This is a flowchart of merging a dependency list of open source software provided by this application;

[0142] Figure 4 It is a flow chart of vulnerability knowledge extraction in a vulnerability identification method provided by this application;

[0143] Figure 5 It is a flowchart of text parsing in a vulnerability identification method provided by this application;

[0144] Figure 6 It is a flow chart of source code parsing in a vulnerability identification method provided by this application;

[0145] Figure 7 It is a schematic diagram of a process of merging and generating a vulnerability knowledge graph in a vulnerability identification method provided by this application;

[0146] Figure 8It is a schematic diagram of a process for parallel generation of an open source dependency graph in a vulnerability identification method provided by this application;

[0147] Fig. 9 This is a schematic diagram of the analysis flow within the open source software process in a vulnerability identification method provided by this application;

[0148] Fig.10 It is a schematic diagram of generating a function call graph in a vulnerability identification method provided by the present application;

[0149] Fig.11A It is a schematic diagram of the inter-process analysis flow of open source software in a vulnerability identification method provided by this application;

[0150] Fig. 11B It is a schematic diagram of function calls between class nodes in the upper layer of a function block in a vulnerability identification method provided by the present application;

[0151] Fig.12 It is a flowchart of converting a call node into a function node in a vulnerability identification method provided by the present application;

[0152] Fig.13 It is a schematic diagram of generating an open source dependency graph in a vulnerability identification method provided by this application;

[0153] Fig.14 This is a flow chart of open source vulnerability call chain path identification in a vulnerability identification method provided by this application. DETAILED DESCRIPTION

[0154] The technical solutions in the embodiments of the present application will be described clearly and in detail below in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0155] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.

[0156] First, an application scenario diagram of an open source vulnerability analysis system provided by an embodiment of the present application is introduced.

[0157] For example, Figure 1 As shown, the application scenario includes an open source software supplier 101, an internal developer 102, an external software user 103, an open source software distribution market 104, compilation and generation 105, an open source vulnerability analysis system 106, a software product distribution market 107 and an alarm assessment module 108, wherein the open source software supplier 101 develops the open source software and then uploads the open source software to the open source software distribution market 104. The internal developer 102 of the software manufacturer searches for suitable open source software in the open source software distribution market according to functional requirements, and relies on it in the self-developed code to compile and generate a complete software product. The compiled software product needs to undergo a security check, and the open source vulnerability analysis system 106 is used to detect whether the open source software relied on in the software product introduces an open source vulnerability. If an open source vulnerability is introduced, the internal developer 102 is notified to rectify it. If there is no risk, the software product is released to the software product distribution market 107 and then downloaded and used by the external software user 103.

[0158] Understandably, Figure 1 For example only, the open source vulnerability analysis system 106 is deployed after the software product is compiled and generated and before it is released. This is the optimal deployment location, but in actual applications, the application scenario may change. It can be applied before compilation and generation, or even in the open source software dependency requirement analysis stage, or deployed after release. It can perform continuous vulnerability scanning on software products and provide real-time feedback on vulnerability warning information. The specifics are not limited here.

[0159] Figure 2 This is an architecture diagram of an open source vulnerability analysis system provided by this application. Figure 2 As shown, the architecture includes open source software dependency list generation S210, open source software dependency list S220, vulnerability knowledge extraction S230, open source dependency graph parallel generation S240, security vulnerability knowledge graph S250 and open source vulnerability warning and risk assessment S260, wherein the number of application software projects that establish a connection relationship with the open source software dependency list generation S210 can be one or more, and the number of application software projects that establish a connection relationship with the open source vulnerability warning and risk assessment S260 can be one or more, and this application does not make specific limitations.

[0160] Each unit of the above-mentioned open source vulnerability analysis system can be deployed on a computing device, wherein the computing device includes a bare metal server (Bare Metal Server, BMS), a virtual machine, a container or an edge computing device. Among them, BMS refers to a general physical server, such as an ARM server or an X86 server; a virtual machine refers to a complete computer system with complete hardware system functions implemented by network function virtualization (NFV) technology and simulated by software, running in a completely isolated environment; a container refers to a group of processes that are resource-constrained and isolated from each other; an edge computing device refers to a device that is closer to the data source and end user and has low latency and high bandwidth characteristics, such as smart routing, edge servers, etc. Each unit of the open source vulnerability analysis system can also be deployed in a computing device cluster, which may include multiple of the above-mentioned computing devices.

[0161] Combination Figure 2 , the functions of each unit module and the vulnerability identification method in the open source vulnerability analysis system 200 are described below.

[0162] S210: The open source software dependency list generation system may at least include step S211, step S212, step S213 and step S214.

[0163] S211: Extract and obtain an open source dependency configuration file from the software product, wherein the open source dependency configuration file records the names of all directly dependent open source software and the introduced version numbers; the software product may include a software application or system, which is developed by one or more developers or teams to implement product functions or provide product services; open source software can be publicly accessed, viewed, used and modified, and needs to comply with a specific open source license agreement; the open source dependency configuration file may be in the form of a text file, which may contain configuration information of the software product and information related to open source dependencies, so the open source dependency configuration file may record the names of the open source software on which the software product depends and the introduced version numbers, and the open source dependency may refer to the open source software components or libraries used by the software product when developing the software product; the direct dependency of open source software may refer to the open source software components directly used by the software product, which are directly associated with the functions and codes of the software product, and the direct dependency of open source software generally includes libraries or modules referenced in the source code of the software product; the introduced open source software version number generally refers to the identifier of the open source software, which is used to indicate a specific version of the software, and the version number may include one or more of the major version number, the minor version number and the revision number, so that developers can understand the updates and changes of the software.

[0164] It should be understood that the open source software dependency list generation system can more easily manage open source dependency configuration files by scanning configuration files, identifying dependencies, and extracting dependency-related information. Extracting open source dependency information from the configuration files of software products and obtaining open source dependency configuration files are important steps in the open source software dependency list generation system.

[0165] S212: Generate a direct dependency list of the open source software according to the name of the open source software and the introduced version number. For example, if the name of the open source software is oss_compent_1 and the version number is 1.0, the dependency list records (oss_compent_1, 1.0);

[0166] It should be understood that in software development projects, other open source software or libraries are generally relied upon to implement specific functions or modules. These dependencies can be divided into direct dependencies and indirect dependencies. Direct dependencies generally refer to open source software directly referenced by the project, and indirect dependencies generally refer to software that is depended on by direct dependencies. The dependency list can list the name and version number of the open source software. Each direct dependency can be identified as a pair of name and version number. The name is the unique identifier of the open source software, which can be the name of the software or the name of the library (for example, oss_compent_1). The version number is a specific version of the open source software, generally represented by one or more combinations of numbers or letters (for example, 1.0). Generating a direct dependency list of open source software can be achieved by parsing the dependency declaration in the configuration file or source code of the application software project, and can be completed through the open source software dependency list generation system or manual operation. The process of parsing the dependency declaration may involve checking which open source software libraries are referenced in the project and their corresponding version numbers.

[0167] Optionally, obtaining a list of open source software is a key step in the design of the open source vulnerability analysis system 200. The list of direct dependencies of open source software may also include detailed information about each dependency, such as the license agreement of the dependent software, which helps to use the open source dependency graph of the real call chain to deduce the propagation path of open source software vulnerabilities between open source software dependency levels.

[0168] S213: According to the open source software names and version numbers recorded in the open source software direct dependency list, download the corresponding version of the open source software from the open source software distribution market, and decompress all the downloaded open source software; the open source software direct dependency list is the list generated in step S212, which can list the names and version numbers of the open source software that the software project depends on. This list can generally be obtained by parsing the project's configuration file (such as package.json, requirements.txt, etc.); the open source software distribution market can be an online platform, from which developers can obtain various open source software. Some common open source software distribution markets include: Python Package Index (PyPI), Node Package Manager (NPM), Docker Hub, and software repositories of various operating systems, etc. These markets usually contain a large number of open source software packages, from which developers can select and download; download and obtain a specific version of the open source software from the open source software distribution market. Developers can use command line tools, package managers, or special download tools to download the required software packages according to the names and version numbers in the open source software list, which can ensure that the downloaded version exactly matches the project's requirements; the downloaded open source software is usually provided in the form of compressed files, such as ZIP, TAR, etc., and the decompression operation decompresses these files into a specified directory of the local system for subsequent installation and use.

[0169] It should be understood that the open source software dependency list generation system can use automated tools (such as package managers, CI / CD tools) to automatically download and decompress open source software, reducing the need for manual operations; in order to increase download speed, local cache or mirroring can be used to reduce frequent visits to the open source software distribution market, which can speed up downloads and ensure system stability; when downloading and using open source software, the software license agreement also needs to be considered to ensure compliance with relevant regulations and compliance requirements.

[0170] S214: Directly dependent open source software will also rely on open source software for development, so it is necessary to loop step S211 for each open source software downloaded from the open source software distribution market, that is, extract the open source dependency configuration file under the open source software directory, and generate an open source software direct dependency list with the open source software name and the introduced version number recorded in the open source software dependency configuration file, which can be called the open source software indirect dependency list of the software product; loop step S214 until no new open source software names and version numbers are added to the open source software indirect dependency list.

[0171] It should be understood that the purpose of parsing complex open source software dependencies is to build a comprehensive open source software dependency list, which may include: one or more of the direct dependencies of the software product (called the open source software direct dependency list) or all indirect dependencies (called the open source software indirect dependency list), which, in the open source vulnerability analysis system 200, helps to determine the source path of the vulnerability in the open source software.

[0172] In a specific implementation, the open source software dependency list generation system will download each open source software in the direct dependency, and then parse their file structure to find the open source dependency configuration file, which may include a list of dependencies of the software; extract the name of the open source software and the introduced version number from these parsed configuration files to form a direct dependency list of the software product; then, the open source software dependency list generation system enters a recursive process (a recursive process is an operation of parsing software dependencies, in which the same operation is repeated in each step until a specific condition is met), for each direct dependency, the system will find whether it has its own dependency configuration file in its open source software directory, if so, the system will repeat the above process to extract the information of the indirect dependency, each indirect dependency performs the same operation, extracts their name and version number, and can form an indirect dependency list of the software product, wherein this process will continue recursively until no new open source software names and version numbers are found, which means that the system has completely explored all dependencies. The open source vulnerability analysis system 200 will further analyze these open source software dependency lists, perform vulnerability detection on each dependency, and find vulnerabilities in the dependencies, and can generate warnings or suggestions to help developers take repair measures, such as upgrading to a secure version or applying patches.

[0173] like Figure 3 This is a flowchart of the open source software dependency list merging process provided by the application. This step can be Figure 2 The open source software dependency list generation system S210 in the embodiment is implemented.

[0174] Step S220: Obtain a list of open source software dependencies of the software product to be analyzed.

[0175] Optionally, the open source software dependency list includes multiple open source software that the software product to be analyzed depends on. The open source software dependency list generation system can merge the direct dependency list and the indirect dependency list of the obtained open source software product into another separate list, which can include the names and introduced version numbers of all open source software that directly and indirectly depend on the open source software; wherein, if there are duplicate items in the summarized open source software dependency list, the system will deduplicate the list and only retain one entry for the same open source software name and version number. Deduplication is to ensure that there are no duplicate dependencies in the final dependency list. Dependencies; During the deduplication process, version numbers require special handling. If the same software name has multiple different versions of dependencies, they should all be retained because different versions of software may have different functions or vulnerabilities. However, for dependencies with the same name and version number, only one is retained. For example, the direct dependency list of a software product has 1 / 2 / 3, and the indirect dependency list has 1 / 2 / 4 / 5 / 6 / 7. When deduplicating, 1 / 2 are duplicated, but 1 is a different version and 2 is the same version, so 2 is merged. Both versions 1.0 and 2.0 of 1 need to be retained and recorded in the open source software dependency list.

[0176] In a specific implementation, the final open source software dependency list may include one or more of all open source software names or version numbers. The open source software dependency list is an important basis for the open source vulnerability analysis system 200 because it can determine the name and version number of the open source software that needs to be checked. The open source vulnerability analysis system 200 can use the final dependency list to detect known vulnerabilities of each open source software. This final dependency list can also be used for continuous monitoring. The system can periodically check the known vulnerability database to see if there is any new vulnerability information.

[0177] The above text introduces the application scenario and open source vulnerability analysis system of the technical solution provided by this application in combination with the figure. Next, Figures 4 to 14 A detailed description of a vulnerability identification method provided by the application, such as Figure 2 As shown, the method comprises the following steps:

[0178] Step S230: Extract vulnerability knowledge. This step can be done by Figure 2 Step S231, step S232, step S233 and step S234 in the embodiment are implemented.

[0179] Optionally, vulnerability information is a detailed description of a software vulnerability, including one or more of the name, description, affected software version, and repair suggestions of the vulnerability. In the embodiment of the present application, "vulnerability knowledge" may include "vulnerability information". The data of vulnerability knowledge mainly comes from different vulnerability databases, which can provide information about known vulnerabilities. Common open source vulnerability databases in the industry include: National Vulnerability Database (NVD), China National Vulnerability Database (CNVD), China National Vulnerability Database (OSV), etc. These databases collect, record and publish vulnerability information worldwide. NVD is a vulnerability database maintained by the National Institute of Standards and Technology (NIST) of the United States. Its main task is to record and publish software vulnerability information worldwide. NVD provides a detailed description of the vulnerability, CVE number assignment, affected software and version, vulnerability severity score (using CVSS), vulnerability disclosure date, etc. NVD data is very important to the global security community and is used for vulnerability management, open source vulnerability analysis and security research; in addition, CNVD is China's National Information Security Vulnerability Database, which is maintained by the China National Computer Network Emergency Response Technical Processing Coordination Center (CNCERT / CC). The main task of CNVD is to record and publish computer system and network security vulnerability information related to China. It is similar to NVD and provides vulnerability descriptions, CVE numbers, affected software and versions, vulnerability severity scores, etc. CNVD is very valuable to security professionals and organizations in China; OSV is a relatively new open source vulnerability database created and maintained by Google. It focuses on recording vulnerability information in open source software. OSV provides detailed information about known vulnerabilities in open source projects, including a description of the vulnerability, CVE number, affected open source projects and versions, and the status of the vulnerability (fixed, not fixed, etc.). The goal of OSV is to help developers, operators, and organizations manage open source software vulnerabilities, especially those that will not come from the official vulnerability database.

[0180] It should be understood that CVE is a Common Vulnerability Identifier System that assigns a unique identifier (CVE number) to each known vulnerability, which helps track and identify vulnerabilities and enables different organizations and tools to share vulnerability information globally; CVSS is a standard method for assessing the severity of vulnerabilities. It assigns a score to each vulnerability to help organizations prioritize vulnerabilities so that resources can be allocated for repair. CVSS takes into account multiple aspects of the vulnerability, including access complexity, the range of affected users, and the extent of the impact; Common Weakness Enumeration (CWE) is a standard method for identifying and recording common software security weaknesses. It helps developers understand potential problems in software and how to prevent these problems. CWE catalogs common types of problems in software vulnerabilities.

[0181] It should be understood that these vulnerability libraries and standardized methods can provide a unified framework for describing, identifying, evaluating and sharing software vulnerability information, which helps security researchers, developers and organizations around the world better understand vulnerability issues and take measures to reduce potential risks. Vulnerability libraries can also provide data sources for open source vulnerability scanning tools and vulnerability management systems to help automate vulnerability identification and processing.

[0182] Figure 4 This is a schematic diagram of the process of extracting vulnerability knowledge in a vulnerability identification method provided by this application, such as Figure 4 As shown, the method includes:

[0183] Step S231: Extract vulnerability knowledge or format. This step can be done by Figure 4 Step S401, step S402, step S403 and step S404 in the embodiment are implemented.

[0184] It should be understood that extracting vulnerability-related information from unstructured text such as open source vulnerability web pages includes various unstructured sources, such as one or more of the open source vulnerability web pages, blog posts, mailing lists, etc., extracting vulnerability-related information and converting it into a structured, easy-to-process form. The purpose of this task is to obtain detailed information about the vulnerability so that it can be included in the open source vulnerability analysis system or vulnerability database for further analysis and management; vulnerability knowledge extraction is an information retrieval process that involves extracting key information from unstructured text. The description of the vulnerability can include one or more of the nature of the vulnerability, the affected software or system, etc. The vulnerability is usually assigned a unique identifier, such as C VE number, that is, vulnerability number, is used to uniquely identify each vulnerability; in unstructured text, it usually contains a lot of irrelevant information, such as advertisements, comments, texts with inconsistent formats, etc. The open source vulnerability analysis system needs to filter out these invalid information to ensure the accuracy of extracted vulnerability information; the extracted vulnerability information needs to be converted into a structured data form, which can be a database or text file. This structured data is in a form that the open source vulnerability analysis system can understand and process, which allows the system to effectively store and retrieve vulnerability information. Therefore, once the information is extracted, filtered and formatted, it can be included in the database of the open source vulnerability analysis system, and the open source vulnerability analysis system can identify, track and evaluate vulnerabilities based on the extracted information.

[0185] Step S401: Obtain open source vulnerability information from the web page of the open source vulnerability.

[0186] It should be understood that the open source vulnerability information obtained from the web page may include various key details of the vulnerability, such as one or more of the vulnerability name, vulnerability description, vulnerability priority, vulnerability destruction method, vulnerability corresponding characteristics, etc. First, the web page data is obtained. The open source vulnerability analysis system obtains web pages containing open source vulnerability information. These web pages usually come from one or more of various vulnerability information websites, community forums, blogs or mailing lists. The open source vulnerability analysis system can use web crawler technology to automatically retrieve these web pages, or receive notifications by subscribing to vulnerability information sources.

[0187] Optionally, once the web page content is obtained, the open source vulnerability analysis system will parse the web page, identify and extract key vulnerability information therein. The vulnerability name is a unique name or identifier of the vulnerability, usually a key tag used to reference the vulnerability; the vulnerability description may detail one or more of the vulnerability's nature, impact, attack scenario, etc.; the vulnerability priority may indicate the importance or urgency of the vulnerability, usually expressed in high, medium, low or numbers; the vulnerability destruction method may describe one or more of the methods or steps of how the vulnerability is exploited to implement an attack; the vulnerability corresponding features may include: features, signs or keywords related to the vulnerability, etc.; other vulnerability information may include any other details related to the vulnerability, such as one or more of the CVE number, affected software or versions, etc. It should be understood that the above examples are for illustration only and are not specifically limited in this application.

[0188] Step S402: Data preprocessing and paragraph and sentence processing.

[0189] It should be understood that the web page is preprocessed, segmented, and sentence-processed to obtain one or more of the description of the open source vulnerability, the hosting address or the vulnerability code snippet. During data preprocessing, the open source vulnerability analysis system needs to preprocess the web page content, which may include removing HTML tags, filtering out non-text content or normalizing the text format, etc.; during segmentation, the web page usually includes different parts, such as title, overview, detail, reference document, etc. The open source vulnerability analysis system divides the web page content into different paragraphs or parts according to these identifiers, and each paragraph usually contains specific information related to the vulnerability; during sentence processing, in each paragraph, the system will further divide the text into sentences. In order to more accurately locate and extract sentences containing vulnerability information, these sentences may include information such as the description, impact, and degree of harm of the vulnerability.

[0190] Optionally, extract open source software vulnerability description sentences. Once the sentences are separated, the system will look for sentences containing key information about the vulnerability, which may include one or more sentences such as the vulnerability name, affected software and version, and the nature of the vulnerability. These sentences constitute the vulnerability description; extract the hosting address. The vulnerability is usually associated with the place where the code is hosted, such as GitHub, GitLab, SourceForge, etc. The system will look for sentences or paragraphs containing these hosting addresses. Once the hosting address is extracted, the system can find the vulnerability code snippet based on these addresses; extract the vulnerability code snippet. The hosting address provides the location of the vulnerability code. The open source vulnerability analysis system can access these addresses to obtain vulnerability-related code snippets. These code snippets can usually include the implementation of the vulnerability, the repair of the vulnerability, or other software components related to the vulnerability. This process allows the open source vulnerability analysis system to extract vulnerability-related information from unstructured web page text and further analyze and utilize this information, which helps the system understand the nature, risk, and potential impact of the vulnerability so that appropriate measures can be taken to deal with the vulnerability problem.

[0191] Step S403: Obtain an open source vulnerability description statement.

[0192] It should be understood that an open source vulnerability description statement is a natural language sentence or paragraph that describes a specific open source vulnerability. It can usually include one or more of the following: vulnerability name, affected version, vulnerability nature, vulnerability priority, other related information, etc. The vulnerability name is the unique identifier or name of the vulnerability, which is used to reference the vulnerability in vulnerability databases and documents; the affected version describes which software versions are affected by the vulnerability, which can usually include the name of the software, version number range, etc.; the vulnerability nature describes the nature of the vulnerability, such as how the vulnerability is exploited, possible attack scenarios, etc.; the vulnerability priority is the importance or urgency of the vulnerability, usually expressed as high, medium, low or a numerical value; the vulnerability description statement can also include one or more other related information such as the vulnerability's CVE number, CVSS score, vulnerability disclosure date, etc.

[0193] Step S404: Open source the vulnerable code snippet.

[0194] It should be understood that open source vulnerability code snippets are code snippets that are affected by the vulnerability in open source software, which can usually include the actual code of the vulnerability and the context of the code snippet, which is very important for developers and security experts because they allow them to gain in-depth understanding of how the vulnerability is implemented and how to fix the vulnerability. Generally, it can include one or more of the following: the specific code of the vulnerability, the context code, or the impact of the vulnerability; the specific code of the vulnerability can be the actual code of the vulnerability, or it can be a function, method, or module, etc.; the context code can be the code around the vulnerability code snippet to help understand how the vulnerability is triggered and how it interacts with other code; the impact of the vulnerability can describe the potential impact of the vulnerability on the system, including possible attacks and vulnerability exploitation scenarios. The open source vulnerability analysis system uses open source vulnerability description statements to identify and classify vulnerabilities, while the open source vulnerability code snippets allow security researchers and developers to gain in-depth understanding of the nature of the vulnerability in order to take appropriate remediation measures. In the system, this information is usually stored in the vulnerability database so that the system can identify and manage known vulnerabilities.

[0195] Figure 5 This is a flowchart of text parsing in a vulnerability identification method provided by this application, such as Figure 5 As shown, the method includes:

[0196] Step S232: Parse the text. This step can be done by Figure 5 Step S501, step S502, step S503, step S504, step S505, step S506 and step S507 in the embodiment are implemented.

[0197] It should be understood that extracting key information from the open source vulnerability description statement may include one or more of the following: the name of the open source software, the affected version number, the affected file, the affected class, the affected function, the affected statement, etc. This process involves the use of multiple natural language processing (NLP) technologies and models to achieve efficient named entity recognition; text parsing may refer to the process of extracting structured information from natural language text. In the open source vulnerability analysis system, the goal of text parsing is to extract key information about the vulnerability from the open source vulnerability description statement so that the system can understand the nature of the vulnerability and the scope of impact; named entity recognition (NER) is a subtask in natural language processing, which is used to identify and classify named entities from text, such as names of people, places, and organizations. In this embodiment, such as the name of the open source software, version number, affected files, affected classes, affected functions, or affected statements, the NER model is a well-trained machine learning model used to automatically identify these named entities in the text; BERT (Bidirectional Encoder Representations from Transformers is a pre-trained deep learning model used to encode text sentences into high-dimensional vector representations. In the context, BERT is used to convert open source vulnerability description sentences into vector representations for further analysis and identification. Bidirectional Long Short-Term Memory (BiLSTM) and Conditional Random Field (CRF) are deep learning models used for sequence labeling tasks to identify and classify named entities in vulnerability description sentences. BiLSTM is a recursive neural network used to capture contextual information in text, while CRF is used to label named entities in the entire sentence. This combination can effectively identify and classify various named entities in open source vulnerability description sentences, such as software names, version numbers, etc. In the open source vulnerability analysis system, this process helps to automatically extract key information about the vulnerability from the vulnerability description sentence, reducing the need for manual processing and improving the efficiency of the system. It can be used to add vulnerabilities to the vulnerability database for further analysis, classification, and repair. In addition, this automation also helps to reduce the time of vulnerability processing, thereby improving the security of software and systems.

[0198] Optionally, labels are defined for named entities, for example, SN: name of open source software; SV: version number of open source software affected by the vulnerability; SFI: source code file in open source software affected by the vulnerability; SCL: class in open source software affected by the vulnerability; SFC: function in open source software affected by the vulnerability; SCO: statement in open source software affected by the vulnerability; O: irrelevant label.

[0199] Step S501: performing word segmentation processing on the open source vulnerability description statement in step S403.

[0200] It should be understood that the open source vulnerability description statement in step S403 is segmented, [CLS] represents the beginning of the sentence, [SEP] represents the punctuation, and the sequence after segmentation is input into the BERT network. BERT includes word encoding, sentence encoding and position encoding.

[0201] Optionally, BERT is a deep learning model for natural language processing tasks. Before using BERT for text encoding, the text usually needs some preprocessing. In this preprocessing, special tags are usually used to indicate the structure and composition of the text. In BERT, [CLS] represents the sentence start tag of "Classification", which is usually placed at the beginning of the text, and [SEP] represents the separation tag between sentences or paragraphs, which is used to tell the model the different parts of the text. These tags can help BERT understand the structure of the text.

[0202] Alternatively, natural language text usually needs to be segmented into words or subwords so that the model can understand the semantics and structure in the text. BERT uses the WordPiece segmentation method to split the text into small units, which can be words or subwords in the vocabulary. These segmented unit sequences are input into the BERT network for encoding.

[0203] Optionally, the BERT model includes multiple nested encoding layers for processing input text, which may include: word encoding, sentence encoding, position encoding, etc.; in word encoding, BERT encodes each unit after word segmentation to capture their semantic information, which can be achieved by mapping the nested representation of word segmentation to a higher-dimensional nested representation; in sentence encoding, BERT is a pre-trained model that is commonly used for natural language understanding tasks, so it has the function of encoding entire sentences or paragraphs in order to understand the context, which is achieved through multi-layer Transformer encoders; in position encoding, BERT also includes position information in the encoding so that the model understands the position of a word or subword in a sentence, which is achieved by encoding the position information into a vector and adding it to the nested representation.

[0204] Referring to the above content, it can be seen that in the open source vulnerability analysis system, BERT can be used to process text information, including vulnerability descriptions, documents, or any text that requires natural language processing. For example, BERT can be used to encode vulnerability description sentences into vector representations for vulnerability classification, association analysis, or other NLP tasks. This helps the system better understand and process text data and improve the accuracy and efficiency of the open source vulnerability analysis system.

[0205] Step S502: Encode each word in the input sequence to obtain a one-hot vector (000001).

[0206] It should be understood that in the preprocessing process of text data, in natural language processing, text data usually needs to be converted into a form that can be understood by computers for further processing and analysis. For each word in the input sequence, word encoding is usually performed to map it into a digital form.

[0207] Alternatively, a common way to encode words is to use one-hot encoding, in which each one-hot vector is a binary vector, and each word is mapped to a unique vector in which only one element is 1 and the other elements are 0. The position of this 1 indicates the index position of the word in the vocabulary. For example, if there is a vocabulary containing 6 words, the word "apple" may be encoded as (0, 0, 0, 0, 0, 1), where the position of "1" indicates the sixth position of "apple" in the vocabulary.

[0208] Optionally, in an open source vulnerability analysis system, text data may include vulnerability descriptions, open source software codes, etc. It is very important to properly encode text data because this enables the system to understand text information and perform tasks such as vulnerability classification, association analysis, and vulnerability identification. Pre-trained deep learning models, such as BERT, allow the system to capture text semantic information in lower-dimensional continuous vectors.

[0209] Step S503: Sentence encoding is performed on the sentence fragment of each word in the input sequence to distinguish different sentences, thereby obtaining a one-hot vector; in natural language processing, sentence encoding refers to determining the sentence or paragraph in which each word or subword is located, and special markers or methods can be used to indicate the boundaries between different sentences or paragraphs, which can help the model distinguish different sentences in the text; the one-hot vector is an encoding method used to represent sentences or paragraphs in a text.

[0210] It should be understood that in an open source vulnerability analysis system, text data usually includes vulnerability descriptions, vulnerability code snippets, software documentation, etc. The purpose of sentence encoding can be to distinguish different sentences, especially when the text contains multiple sentences or paragraphs, which can help the system understand the structure of the text, especially when there are multiple sentences or paragraphs in the text related to the vulnerability. In an open source vulnerability analysis system, sentence encoding can be used to mark which sentences contain vulnerability descriptions and which sentences contain code examples, so that the system can better understand and process vulnerability information.

[0211] Step S504: Perform position encoding on the position of each word in the input sequence to distinguish the context information of the input sequence and obtain a one-hot vector (000001).

[0212] It should be understood that in natural language processing, the order and position of words usually have an important impact on the meaning of a sentence or paragraph. The position of each word in the text is encoded to distinguish contextual information, and it is mentioned that different encoding methods can be used, such as one-hot vector, word2vec, GloVe, etc. Word2vec and GloVe are word embedding technologies, which embed words into a continuous vector space to capture the semantic relationship between words.

[0213] Step S505: sum the word codes, sentence codes and position codes obtained in steps S502, S503 and S504.

[0214]

[0215] Among them, O tok , O seg , O pos Represents the one-hot vectors of word encoding, sentence encoding and position encoding respectively, W tok , W seg , W pos Represent the parameter matrices of word encoding, sentence encoding and position encoding respectively, |V|, |S|, |P| represent the word vector dimension, the number of sequence sentences, and the maximum number of positions respectively, and H is the embedding dimension.

[0216] Step S506: the summed and merged vector code Eword is input into the BiLSTM and CRF models for processing. The BiLSTM can capture the association information between the sequence data, and the CRF predicts the association information extracted by the BiLSTM and extracts the predefined labels.

[0217] It should be understood that the Eword vector is a vector representation of processed text data. This vector contains information about each word or subword in the text. The purpose of the encoding process is to convert text data into a form that can be processed by a computer for subsequent natural language processing tasks. The vulnerability description text may contain various information, such as the affected software version, vulnerability type, etc. By converting the text into sequence data and then using the BiLSTM and CRF models, the system can automatically identify and mark important information in the text to better understand and process the vulnerability description, which helps to improve the accuracy and efficiency of the open source vulnerability analysis system and the automated analysis of vulnerabilities.

[0218] Step S507: According to the output result of step S506, named entity information such as the open source software name, affected version number, affected file, affected class, affected function, affected statement, etc. can be extracted from the open source vulnerability description statement.

[0219] From the above description, it can be seen that after extracting vulnerability knowledge from unstructured text in the open source vulnerability web page based on the open source vulnerability library, it also includes extracting the open source vulnerability description statement of the vulnerable open source software from the vulnerability knowledge, and performing text parsing on the open source vulnerability description statement to obtain vulnerability named entity information. The vulnerability named entity information includes the name information, version information, vulnerability statement and at least one of the following items of the vulnerable open source software: vulnerability file information, vulnerability class, vulnerability function. In the process of generating the vulnerability knowledge graph of the vulnerable open source software, the source code information extracted by analyzing the open source vulnerability source code fragment is integrated with the vulnerability named entity information.

[0220] In addition, text parsing is performed on the open source vulnerability description sentences to obtain vulnerability named entity information. This process includes word segmentation of the open source vulnerability description sentences to obtain a sequence after word segmentation, and processing the sequence after word segmentation through a Transformer-based bidirectional encoder representation network to obtain word encoding, sentence encoding, and position encoding, which are summed and merged. Finally, the merged vector encoding is processed through a bidirectional long short-term memory model and conditional random field to output the vulnerability named entity information.

[0221] Figure 6 This is a schematic diagram of the source code parsing process in a vulnerability identification method provided by this application. Figure 2 The method shown includes step S233: source code parsing, which can be performed by Figure 6 Step S601, step S602, step S603, step S604, step S605 and step S606 in the embodiment are implemented.

[0222] It should be understood that the open source vulnerability source code fragment extracted in step S231 is analyzed to extract the source code information such as the files, classes, functions, and statements affected by the open source vulnerability. The source code information is defined as SFI: source code files affected by the vulnerability in the open source software; SCL: classes affected by the vulnerability in the open source software; SFC: functions affected by the vulnerability in the open source software; and SCO: statements affected by the vulnerability in the open source software.

[0223] Step S601: Code slicing: forward slicing the open source vulnerability code snippet.

[0224] It should be understood that code slicing is the process of breaking down the code in open source software into smaller code segments for more careful analysis and processing. Forward slicing means starting from the starting point of the code (usually the entry point of the program), gradually selecting the execution path, and generating code snippets related to the execution path. This is used to analyze the execution flow of the program, which can help the system determine how the vulnerability is triggered and understand the root cause of the vulnerability.

[0225] Step S602: Generate an abstract syntax tree, and parse the source code fragment into syntax nodes of the abstract syntax tree.

[0226] It should be understood that an abstract syntax tree (AST) is a tree data structure used to represent the grammatical structure of source code. The source code can be parsed into a tree structure, in which each node represents a grammatical element in the source code, such as a variable, function, operator, etc. The generation of AST is part of a compiler or interpreter, which helps to understand and operate the source code. In order to build AST, source code fragments need to be parsed. Parsing is the process of converting source code text into a data structure that a computer can understand and process. Different programming languages ​​have different parsers and parsing rules. In Python, the ast.parse() function in the built-in ast module is usually used to parse Python code.

[0227] Step S603: Map the label to the abstract syntax tree, and add a label node to the node in the abstract syntax tree through structured caching.

[0228] It should be understood that "labels" can generally mark or classify specific grammatical structures or patterns in the source code. These labels can represent the nature of the code, such as the vulnerability type (e.g. SFI-Secure Function Injection, SCL-Security Configuration Vulnerability), etc. Mapping to the abstract syntax tree means associating these labels with the corresponding nodes in the AST. Structured cache generally refers to a data structure used to store and organize data to improve access speed and efficiency. Structured cache is used to store label information related to AST nodes for fast access and retrieval.

[0229] Optionally, by associating tags with AST nodes, the open source vulnerability analysis system can identify specific vulnerability patterns or code structures, making it easier to identify and classify potential vulnerabilities. For example, SFI and SCL tags may be used to indicate code snippets associated with function injection or configuration vulnerabilities. The system can generate a vulnerability report based on the tag information, which contains the nature, location, and recommended fixes of the vulnerability, which helps developers more easily understand the nature of the vulnerability and take measures to resolve the problem. Tags can be used for automated analysis to evaluate potential vulnerabilities in the code. The system can execute rules and policies based on tag information to improve the accuracy and efficiency of vulnerability identification.

[0230] Step S604: Map the instance in the code snippet to the abstract syntax tree, and add an instance node to the node in the abstract syntax tree through structured caching.

[0231] It should be understood that in source code parsing, "instance" can generally refer to specific variables, classes, functions or objects in the source code, etc. Mapping these instances to corresponding nodes in AST means associating entities with their locations in the code so that the system can understand how various entities in the code are related to each other, and structured cache is used to store instance information related to AST nodes for fast access and retrieval, such as shared_data.py, SharedDataMiddleware and other instances.

[0232] Optionally, vulnerability propagation chains often involve interactions between multiple code entities. By mapping instances to AST, the system can more easily identify potential vulnerability propagation chains. For example, if a variable is defined in one function and then used in another function, the system can track this propagation through instance mapping. In the propagation chain, there may be sources and victim nodes involved in the vulnerability. Through instance mapping, the system can associate specific code entities to determine how the vulnerability propagates from the source to potential victim nodes, which helps to identify key vulnerability propagation chains.

[0233] Step S605: According to the statement with open source vulnerability marked by “#code”, find the syntax node with label node SCO and instance node code in the abstract syntax tree, which is called leaf node.

[0234] It should be understood that "#code" can be a comment or mark used in the source code to indicate a specific part of the code, and is used to identify statements of potential open source vulnerabilities. The mark is a clue to locate the problematic code segment. AST is a structured representation of the source code, which presents the grammatical structure of the source code in a tree structure. AST usually includes various grammatical nodes.

[0235] Optionally, the system scans the source code for comments or tags with a #code tag. Then, the source code is parsed to build an AST. The system traverses the AST to find syntax nodes with a label node of SCO and an instance node of code. These nodes are considered to be code snippets related to potential vulnerabilities.

[0236] In the specific implementation, the #code tag identifies the vulnerability, and the system can track the vulnerability propagation chain starting from this node to determine how the vulnerability propagates. Finding a specific leaf node helps to identify the specific code segment related to the vulnerability. By associating comments or tags with AST, the vulnerability can be located, which helps to analyze the vulnerability propagation path.

[0237] Step S606: Perform path traversal through leaf nodes, use the path sorting algorithm to deduce the relationship between syntax nodes, and generate a propagation path of the open source vulnerability within the open source software.

[0238] It should be understood that "leaf node" usually refers to the terminal node in AST, which can represent the statement in the source code. Path traversal is a method of traversing AST to find the relationship between specific nodes, which can be implemented by graph search algorithms, such as depth-first search (DFS), breadth-first search (BFS), minimum spanning tree algorithm and other traversal algorithms; path sorting algorithm is used to deduce the relationship between syntax nodes, especially the vulnerability propagation path in the source code, which helps the system determine the correlation between code segments in order to generate vulnerability propagation path; vulnerability propagation path is represented by quadruple, each quadruple includes four values: file name (file), class name (class), function name (function) and statement (statement). Class name and function name can have multiple values, represented by (function1:function2), and two propagation paths of open source vulnerabilities in open source software can be obtained: (shared_data.py:SharedDataMiddleware: (get_package_loader:loader):code), (shared_data.py:SharedDataMiddleware: (get_directory_loader:loader):code).

[0239] Optionally, the system traverses the AST, starting from the leaf nodes, to find statements containing the #code tag, and then analyzes the relationship between these leaf nodes. The path sorting algorithm is used to deduce the relationship between syntax nodes to generate a vulnerability propagation path, which helps to intuitively represent how the vulnerability propagates from one node to another.

[0240] In the specific implementation, generating a quadruple representation of the vulnerability propagation path helps the system visualize how the vulnerability propagates in the code and helps to better understand the vulnerability propagation mechanism. Therefore, the system can further analyze the generated path and identify potential vulnerability propagation chains.

[0241] From the above description, we can see that this process is to slice the open source vulnerability source code fragment, parse the source code fragment into the syntax node of the abstract syntax tree, add node labels to the nodes of the abstract syntax tree, map the source code fragment to the abstract syntax tree, add instance nodes to the abstract syntax tree, and then traverse the path of the leaf nodes of the abstract syntax tree to generate a propagation path from the root node to the leaf node. The leaf node is a label node marked with the open source vulnerability.

[0242] Figure 7 This is a schematic diagram of the process of merging and generating a vulnerability knowledge graph in a vulnerability identification method provided by this application, for Figure 2 The method shown includes step S234 of storing vulnerability information, which may specifically include:

[0243] It should be understood that the information extracted in step S507 and step S606 are merged to obtain complete vulnerability information, including: vulnerability information, open source software name, affected version number, affected files, affected classes, affected functions, affected statements and other information.

[0244] It should be understood that the vulnerability knowledge graph is stored in a graph database such as neo4j; has_*=1 represents the existence of an edge, has_*=0 represents the non-existence of an edge. In the vulnerability knowledge graph, the software name, affected version, and affected statement must exist, otherwise it is an invalid vulnerability knowledge graph and needs to be rebuilt or discarded.

[0245] Optionally, in order to effectively manage vulnerability information, this information can be stored in a graph database, such as Neo4j. Neo4j is a graph database that stores data in the form of nodes and relationships. It is suitable for storing and querying data with complex relationships, such as knowledge graphs, network security and vulnerability management. Vulnerability information usually involves the relationship between multiple entities, such as software name, version, file, class, function, etc. In a graph database, nodes represent entities (such as software name, version, file, etc.), and edges represent the relationship between different entities. For example, has_*=1 can indicate the existence of a relationship, such as the association between a file node and an affected function node, and has_*=0 can indicate the non-existence of a relationship, such as the lack of association between an affected version node and a software name node.

[0246] In the specific implementation, the vulnerability knowledge graph based on the graph database can be used to build a vulnerability propagation chain. By traversing the associated nodes and edges, it is possible to identify how the vulnerability propagates from one entity to another, for example, from a software version to a file, then to a class and a function, and finally affecting a certain statement; through the graph database, complex path queries can be performed to find the propagation path of a specific vulnerability.

[0247] Figure 8 This is a flow chart of parallel generation of an open source dependency graph in a vulnerability identification method provided by this application, such as Figure 8 As shown, the method includes:

[0248] against Figure 2 The method shown includes step S240: parallel generation of an open source dependency graph.

[0249] This step can be done by Figure 8Step S801, step S802, step S803, step S804, step S805, step S806, step S807, step S808, step S809, step S810, step S811 and step S812 in the embodiment are implemented.

[0250] It should be understood that the generation of the open source dependency graph and the extraction of vulnerability knowledge are synchronous behaviors. For the open source dependency graph, due to the explosive number of open source dependent software and the complex dependencies between the levels of open source dependent software, the present invention uses parallel computing between subgraphs to generate it, first generating subgraphs and then merging the subgraphs to generate the open source dependency graph.

[0251] Optionally, before carrying out the specific steps, the open source software is defined first: a certain software product P has an open source software dependency list P = a set of {SN1, SN2, ..., SN3}, where SN* is the name of the open source software; each open source software SN* = a set of {SFI1, SFI2, ..., SFIn}, where SFI& is a source code file; each source code file SFI& = a set of {SCL1, SCL2, ..., SCLn}, where SCL# is a class; each class SCL# = a set of {SFC1, SFC2, ..., SFCn}, where SFC@ is a function; each function SFC@ = a set of {SCO1, SCO2, ..., SCOn}, where SCO% is a statement, and a statement is the smallest unit in the open source dependency graph generation algorithm.

[0252] Step S801: Generate a corresponding number of process pools according to the number of open source software items in the open source software dependency list.

[0253] It should be understood that the process pool, including the name and version number of the open source software, is a parallel computing mechanism used to distribute tasks among multiple processes and execute them simultaneously. Each process runs in an independent execution environment and can handle different tasks. Task parallelism can be achieved by generating processes equal to the number of open source software entries in the open source software dependency list, and each process can independently handle an open source software dependency.

[0254] Step S802: Each open source software occupies a process pool for analysis, and the number of source code files in the open source software is counted to generate a corresponding thread pool. One process pool corresponds to multiple thread pools.

[0255] It should be understood that a process pool is a group of processes used to execute tasks in parallel. Each open source software is assigned to an independent process for parallel analysis. Each process runs independently and has its own memory and execution environment. For each open source software, the number of its source code files needs to be counted, which helps determine how many threads need to be allocated to process the source code files. Different software may contain different numbers of source code files. The number of thread pools is related to the number of source code files of each open source software. One process corresponds to one open source software, and there are multiple threads in a process, each thread is responsible for processing one source code file.

[0256] Step S803: Analyze each source code file.

[0257] It should be understood that the source code file contains the source code of the software program, usually in the form of a text file. Each source code file usually includes one or more code units, such as functions, classes, variable definitions, etc. The source code file is studied in depth to understand its structure, syntax, and functionality.

[0258] Step S804: Each file contains 0 to several classes {SCL1, SCL2, ..., SCLn}, and the file is forward sliced ​​to obtain sub-code fragments of several classes.

[0259] It should be understood that source code files are source code files in a software project, which may contain different numbers of classes or code segments, with curly brackets used to represent a collection of these classes, where SCL1 represents the first class, SCL2 represents the second class, and so on, until SCLn represents the nth class; forward slicing is a source code analysis method, which starts from the starting position of the source code file, and gradually extracts code segments according to the code structure and syntax until a termination condition is encountered, and forward slicing is applied to each source code file to obtain several code fragments in the file; each code fragment is a subset in the source code file, which usually contains code related to one or more classes, therefore, from the forward slicing of each source code file, multiple sub-code fragments can be obtained, each of which is related to a class in the source file.

[0260] Step S805: Each class sub-code fragment contains a number of functions {SFC1, SFC2, ..., SFCn}, and each class is forward sliced ​​to obtain sub-code fragments of a number of functions.

[0261] It should be understood that the above-mentioned forward slicing of the source code file obtains the sub-code fragments of each class, and each class sub-code fragment is further subdivided into several functions, where SFC1 represents the first function in the class sub-code fragment, SFC2 represents the second function, and so on, until SFCn represents the nth function, and forward slicing is performed on each class sub-code fragment, starting from the starting position of each class sub-code fragment, and gradually extracting the sub-code fragments of the function according to the code structure and grammatical rules; after forward slicing each class sub-code fragment, the sub-code fragments of several functions contained in the class can be obtained, and the sub-code fragment of each function generally includes the code content of the function.

[0262] Step S806: Each function sub-code fragment contains a number of statements {SCO1, SCO2, ..., SCOn}, and each function is forward sliced ​​to obtain a number of statements.

[0263] It should be understood that the above-mentioned sub-code fragments of each class are forward sliced ​​to obtain the sub-code fragments of each function, and each function sub-code fragment is further subdivided into several statements, where SCO1 represents the first statement in the function sub-code fragment, SCO2 represents the second statement, and so on, until SCOn represents the nth statement; each function sub-code fragment is forward sliced, starting from the starting position of each function sub-code fragment, and gradually extracting the statements in the function according to the code structure and grammatical rules; after forward slicing each function sub-code fragment, we obtain several statements contained in the function, each statement usually represents a logical operation or command in the code, and they are combined together to constitute the function of the function.

[0264] Step S807: Treat each statement as a node and generate nodes for all statements.

[0265] It should be understood that in the source code, each line or code block represents a statement. These statements can be various operations, conditions, loops, etc. in the program. Each statement is regarded as an independent node. A node is created for each statement in the source code, and these nodes are combined into a graph structure, in which each node represents a statement.

[0266] In summary, after obtaining the open source software dependency list for the software product to be analyzed, the open source software dependency list may include multiple open source software that the software product to be analyzed depends on. Therefore, the process of generating the function call graph of each open source software in the open source software dependency list is as follows: first, for each open source software in the open source software dependency list, each open source code file in the open source software is decomposed; secondly, for each open source code file in the open source software, each class in the open source code file is decomposed; again, for the code snippet of each class in the open source code file, each function in the class is decomposed; finally, for the code snippet of each function, the statements contained in the function are decomposed.

[0267] The purpose of this process is to decompose the structure of each open source software into smaller, analyzable units in order to better understand its internal dependencies and the calling relationships between functions. By following this hierarchical decomposition method, a function call graph of each open source software can be constructed, which includes the calling relationships between functions, so as to better analyze and evaluate the security of the software product to be analyzed in subsequent vulnerability analysis and risk assessment.

[0268] Step S808: For statements in the same function slice, a control flow graph (CFG) is generated according to an intra-procedural analysis algorithm.

[0269] It should be understood that for each statement within the same function, a control flow graph is generated with statements as nodes, which may include dividing the function by statements, with each statement as a statement node; if the statement has the behavior of calling other functions, the call node is determined, and the call node is saved based on a structured cache, wherein a control flow graph is generated based on statement nodes and call nodes, and the edges between nodes in the control flow graph represent the execution order of the nodes or the call relationship between the nodes.

[0270] It should be noted that, first, the code inside the function is divided into statements, which means dividing the function into multiple basic blocks, each of which contains one or more statements, and each statement will become a node in the control flow graph. Inside the function, if a statement calls other functions, then the statement will be marked as a call node. This is to record which statements trigger the execution of other functions. Then, the information of the call node will be stored in a structured cache so that it can be accessed and identified in subsequent analysis which statements triggered the function call, and a control flow graph is generated based on the statement nodes and call nodes.

[0271] It should be understood that in the control flow graph, a connection is established between the statement node (representing the statement in the code) and the call node previously saved in the structured cache, that is, there is an edge, which represents the control flow path from a statement node to a call node, which means that when this statement node is executed, the program will jump to the call node saved in the structured cache to execute the corresponding function.

[0272] Optionally, a function usually contains multiple statements, which are executed in a specific logical order. These statements may include conditional statements, loops, assignment statements, etc. A control flow analysis algorithm can be applied to analyze the execution flow of statements within the function to determine the control relationship between each statement. This algorithm will consider conditional branches, loop structures, function calls, etc. to construct a control flow graph. A control flow graph is a graphical representation used to describe the execution flow of a program. It consists of nodes (representing statements or code blocks) and directed edges (representing the direction of control flow). The edges between nodes represent the execution path of the code.

[0273] Step S809: For functions in the same class slice, a function call graph (FCG) is generated according to an inter-procedural analysis algorithm.

[0274] It should be understood that the function call graph within the same class is generated based on the control flow graph, all the control flow graphs CFG of step S808 are merged, and the calling relationship between functions is deduced to obtain the dependency relationship of all function call graphs FCG; in a code snippet of a class (or object), multiple functions are usually included, and these functions represent different methods or operations within the class. Through the inter-function analysis algorithm, the calling relationship between these functions can be analyzed; the function call graph indicates which function calls which function, and the relationship between them, and the calling path between functions can be tracked; the function call graph is usually a directed graph, in which nodes represent functions, and directed edges represent the calling relationship between functions. This process helps to understand the relationship and calling method of functions within the class, and the generated function call graph can be used to analyze the execution flow, dependency, etc. of the program.

[0275] In the specific implementation, for example, there is a class A, which contains functions func1, func2 and func3. The step is to extract the functions func1, func2 and func3 in class A and the calling relationship between them from the entire control flow graph, which will form a graph called "function call graph within the same class", in which nodes represent these functions and edges represent the calling relationship between them, which helps to analyze the interaction and calling relationship between functions within a specific class.

[0276] Step S810: For all classes and functions in the same file, all function call graphs FCG of step S809 are merged, and the calling relationships between classes and functions are deduced to obtain the dependency relationships of all function call graphs FCG, thereby generating the function call graph FCG of the entire file.

[0277] It should be understood that the function call graphs of all classes in the same open source code file are merged to generate the function call graph of the open source code file. It can be understood that an open source code file usually contains definitions of multiple classes, and these classes can contain various functions to implement different functions. The goal of this step is to merge the function call graphs of all classes in the same open source code file into a large function call graph. For example, there is an open source code file named "example.py", which contains two classes: Class A and Class B. Class A contains functions funcA1 and funcA2, and Class B contains functions funcB1 and funcB2. For each class, a function call graph can be generated to represent the calling relationship between functions in this class. These graphs will include nodes (functions) and edges (calling relationships between functions). Finally, the two in-class function call graphs are merged into a large function call graph, which includes all classes and functions in the entire "example.py" file, so as to comprehensively analyze the code interactions and functions of the entire file.

[0278] Optionally, a software file (usually a source code file) may contain multiple classes and functions, which represent different functions and modules within the file. In step S809, a function call graph FCG has been generated for each class, and the function call graphs FCG are merged together to represent the calling relationship between functions; the merged function call graph is analyzed to determine the relationships and dependencies between classes and functions, which may include which function calls which function, which class calls which function, and the calling paths and passed parameters between them; by analyzing the merged function call graph, the dependencies between different functions, including direct and indirect dependencies, can be obtained, and finally a function call graph for the entire file is generated, which represents the calling relationship between all functions in the file, including the interactions between classes and functions, and their dependencies.

[0279] Step S811: For all files in the same open source software, the function call graph FCG of step S810 is merged to form the function call graph FCG of the entire open source software. The function call graph FCG generated in this step is also called a subgraph of the open source dependency graph.

[0280] It should be understood that the function call graphs of each open source code file in the same open source software are merged to generate the function call graphs of each open source software in the open source software dependency list. Open source software is usually composed of multiple source code files, each file contains different functions and classes. For the files in the same open source software project, in step S810, a function call graph FCG has been generated for each file, which represents the calling relationship between the functions in the file. In this step, these separate function call graphs need to be merged together to form the function call graph of the entire open source software project; the merged function call graph represents the calling relationship between the functions in the entire open source software project, and shows the interaction between the functions in different files; in this process, the generated function call graph can be called a subgraph of the open source dependency graph.

[0281] For example, consider an open source software library in Python, such as "requests". This library is used to process HTTP requests. In this library, there are multiple source code files, each of which contains some functions or modules, such as processing HTTP requests, processing Cookies, etc. To understand how the entire "requests" library calls internal functions and modules, it is necessary to merge the function call graphs in each source code file into a large function call graph. Suppose there are two source code files: request_core.py and cookie_handler.py, each of which contains some functions. A function call graph can be generated for each file to show the calling relationship between functions. For example, the function in request_core.py may call the function in cookie_handler.py. Then, by merging these two function call graphs together, the function call graph of the entire "requests" library can be obtained, which includes the functions in all source code files and the calling relationships between them.

[0282] Step S812: merge the subgraphs of all open source software in the open source software dependency list, and deduce the calling relationships between the open source software to obtain the dependency relationships of all subgraphs, thereby generating an open source dependency graph.

[0283] It should be understood that the open source software dependency list records the list of other open source software that the open source software project depends on. Each open source software project may depend on multiple other projects, and these dependencies may involve different versions and levels. For each open source software project in the list, a subgraph can be generated. The generated subgraph represents the calling relationship between functions in the project, and each subgraph represents the internal structure of an open source software project. In this step, all subgraphs need to be merged together, and the subgraphs of different projects need to be integrated into a large graph to represent the entire open source software dependency graph. The merged graph includes the function calling relationships within different open source software projects. In this step, the calling relationships in the graph need to be analyzed to understand the interactions and dependencies between different projects, which may include which project calls the function of which project, as well as the calling paths and passed parameters between them. Finally, an open source dependency graph is generated, which represents the dependencies and interactions between different open source software projects. This graph is a large-scale structure that represents the entire open source software ecosystem, which helps to understand the interdependence between projects.

[0284] against Figure 2 Step S241 in the method shown: analysis within the open source software process may involve steps S807 and S808, which are used to generate statement nodes and a control flow graph CFG. The control flow graph is an abstract representation of the source code, representing all paths that will be traversed during the execution of a code, and represents the possible flow of all basic block statements in a function in the form of a graph.

[0285] Combination Figure 8 , Step S807 and Step S808: Involving the analysis within the open source software process, this step can be Fig. 9 Step S901, step S902, step S903 and step S904 in the embodiment are implemented.

[0286] Fig. 9 This is a schematic diagram of the analysis process of the open source software process in a vulnerability identification method provided by this application, such as Fig. 9 As shown, the method includes:

[0287] Step S901: Divide the function code fragment into separate basic blocks according to statements, with each statement being a node.

[0288] It should be understood that the function code snippet will be divided into multiple small units, which are determined by the boundaries of statements. Each statement is usually separated by semicolons, line breaks or other grammatical rules. This process divides the code snippet into different statements; basic blocks are the smallest execution units consisting of a single statement, and each statement is regarded as an independent basic block. These basic blocks are usually the basic units of program analysis and optimization, and they can help understand the execution flow and control structure of the code; in this process, each basic block is represented as a node, which means that in the generated data structure, each basic block has a corresponding node, and there are connections between the nodes to represent the control flow relationship between them. This process subdivides the code within the function into smaller, easier-to-analyze units, and the representation of each statement as a node helps to build control flow graphs, data flow graphs, and other data structures related to code analysis.

[0289] Step S902: Each statement node is divided into a non-call node and a call node according to whether there is a behavior of calling other functions. For the call node, a structured cache needs to be used to save a call node. There is an edge between the statement node and the call node.

[0290] It should be understood that each statement is represented as a node, and these nodes constitute a set of nodes in the entire program or function; non-call nodes refer to statement nodes that do not contain call behaviors to other functions. These nodes represent ordinary statements in the program, which can be assignments, conditional statements, loops, etc.; call nodes refer to statement nodes that contain call behaviors to other functions. These nodes represent calls to functions in the program, and the execution flow may be transferred to other functions for execution; structured cache is a data structure used to store information related to call nodes. This information may include the name, parameters, return values, etc. of the called function. The structured cache is used to record and manage the information of call nodes for subsequent analysis and processing; in this process, an edge is created to connect the relationship between call nodes and statement nodes. This edge represents the connection between the call node and the called function, which can usually include the name and parameter information of the called function.

[0291] Optionally, when obtaining whether there is a call node in step S902, hidden nodes need to be considered to correctly display the call relationship, for example: a=b; c=a; d=c; pass=d.word, the real call node of pass is b.word, not d.word. To restore the hidden node to the real call node, the function code fragment needs to be abstracted to generate an abstract syntax tree, and the abstract syntax tree is traversed upward with the hidden node as the leaf node to find the root node, that is, the real call node.

[0292] It should be understood that hidden nodes refer to function call relationships that are not explicitly displayed in the code and are indirectly transmitted through intermediate variables or other means. Direct function calls cannot be seen in the code, but there are actually connections between functions. In the abstract syntax tree, the structure and hierarchical relationship of the code can be reflected. In order to find the real call relationship of the hidden node, it is necessary to traverse the abstract syntax tree. Specifically, the abstract syntax tree can be traversed upward from the hidden node (such as d.word) as a leaf node to find the root node, which represents the real call relationship (such as b.word). This process can solve the problem of hidden nodes so as to accurately restore the call relationship.

[0293] Step S903: If there is a call node in step S902, then step S903 needs to be executed to perform data dependency analysis. Data dependency analysis uses a context-sensitive algorithm, that is, different method call contexts make different analyses for calls with different parameters, and the same function is analyzed multiple times. For example, in the function code snippet, an external function foo is called, where a and b both call the foo function and pass in different parameters x and y. In step S902, two call nodes are obtained.

[0294] In a specific situation, pseudo code can be used to illustrate:

[0295] def foo(param):

[0296] pass#Definition of function foo

[0297] def some_function():

[0298] a=foo(x)

[0299] b = foo(y)

[0300] It should be understood that some_function calls function foo twice, passing parameters x and y respectively, which will result in two different call nodes being generated in step S902; in this case, data dependency analysis is required to understand how different calling contexts affect program behavior, where the algorithm considers different calling contexts, i.e., parameters x and y, and performs analysis for each case, which helps to identify the impact of function calls and how data is passed between functions and affects program behavior.

[0301] From the above description, it can be seen that the calling node is determined according to step S902 and step S903. If there is a hidden node, the hidden node is traversed upward as a leaf node, and the root node found is used as the calling node; if there is a calling node of step S902, different data dependency analysis is required for calls with different parameters to determine the calling node.

[0302] Step S904: Generate a control flow graph CFG for the basic blocks in the function.

[0303] Specifically, combined Fig.10 A schematic diagram of generating a function call graph in a vulnerability identification method provided by the present application is shown. Each node in the control flow graph CFG is a basic block. The child nodes of the node may have call nodes. The function code main() has 6 basic block statements {SCO1, SCO2, SCO3, SCO4, SCO5, SCO6}, among which SCO3 and SCO5 have call nodes, and the edges represent the flow direction of the basic blocks.

[0304] It should be understood that the main() function contains 6 basic block statements, marked as SCO1, SCO2, SCO3, SCO4, SCO5 and SCO6. These basic blocks represent the execution flow of the program, and each basic block contains a set of consecutive statements; there are call nodes in basic blocks SCO3 and SCO5, indicating that these two basic blocks contain calls to other functions, and the call nodes usually represent the starting point of the function call; the edges in the control flow graph represent the flow relationship between the basic blocks, that is, the order of program execution. These edges describe the control flow between the basic blocks and show the conditional branches, loops and execution order of statements in the program.

[0305] against Figure 2 Step S242 in the method shown: open source software inter-process analysis, involves steps S809, S810 and S811, which are used to generate a function call graph FCG, which represents the calling relationship between functions in the entire class and source code file. It can be understood as generating a function call graph for each open source software in the open source software dependency list, where the nodes are function methods and the edges represent calling relationships.

[0306] Combination Figure 8 Step S809, step S810 and step S811: involve open source software inter-process analysis, which can be done by Fig.11A Step S111, step S112, step S113 and step S114 in the embodiment are implemented.

[0307] Fig.11A This is a schematic diagram of the inter-process analysis flow of open source software in a vulnerability identification method provided by this application, such as Fig.11A As shown, the method includes:

[0308] S111: Obtain all function blocks in the class and source code files, generate an abstract syntax tree, find all class nodes and function nodes, and then search down for function nodes in the class nodes to obtain all function blocks.

[0309] It should be understood that the source code file is parsed to generate a corresponding abstract syntax tree. The abstract syntax tree is a tree-like data structure that represents the structure of the source code, including classes, functions, statements, etc. Class and function nodes can be found in the abstract syntax tree. Class nodes usually represent class definitions in source code files, and function nodes represent function definitions in source code files. These nodes are usually represented by specific identifiers or structures in the abstract syntax tree. In the class node, you can continue to search downward for the function node, because a class usually contains multiple methods (functions), so you need to find the function node inside the class node. Once the function node is found, you can further extract the function block. The function block is a section of code within the function, including the function definition and the statements therein.

[0310] Step S112: All function block names are stored in the cache.

[0311] It should be understood that during the code analysis process, all function block names (usually identifiers or names of functions) are stored in a cache. The purpose of this cache is to quickly access and retrieve the names of these function blocks in subsequent operations without having to parse the source code file again to obtain the names of the function blocks. By storing the function block names in the cache, it is possible to avoid re-parsing the same source code file each time the analysis is performed, thereby reducing redundant work.

[0312] Fig. 11B This is a schematic diagram of function calls between class nodes in the upper layer of a function block in a vulnerability identification method provided by this application, such as Fig. 11B As shown, the method includes:

[0313] Step S113: If there is a class node on the upper layer of the function block, the class node is the root node and the function block is a leaf node. There must be an edge between the root node and the leaf node. If there is a calling relationship between the leaf nodes SFCi and SFCj, there is a calling edge between SFCi and SFCj, pointing from node SFCi to node SFCj. If the function called by the leaf node SFCi is not in the class, it is recorded as a calling node, and the calling node name is the function name stored in step S112.

[0314] It should be understood that based on the control flow graph, there must be an edge between the root node and the leaf node, which can be understood as generating an edge between the first root node and the first leaf node; if there is a calling relationship between the leaf nodes SFCi and SFCj, then there is a calling edge between SFCi and SFCj, which can be understood as generating an edge between the second leaf node and the third leaf node.

[0315] It should be understood that the above description of the relationship between function blocks and their relationship with class nodes. If there are class nodes and function blocks in the code, where the class nodes are located at the upper layer and the function blocks are located at the lower layer, just like the class contains methods. In this case, the class node is regarded as the root node and the function block is regarded as the leaf node; among them, there must be an edge pointing from the root node (class node) to the leaf node (function block), that is, the function block with the class node in the upper layer is the first leaf node, the class node with the function block in the lower layer is the first root node, the function block that calls other functions is the second leaf node, and the function block called by other functions is the third leaf node.

[0316] For example, a class Car1 contains a function A, then in the control flow graph, there will be an edge between function A (the first leaf node) and the Car class (the first root node); if function block B (the second leaf node) calls other function blocks C, there will also be edges between these called function blocks C (the third leaf node), then in the control flow graph, there will be an edge from function C to function B; if a function is not included in the class and calls a function in class Car2, then in the control flow graph, an edge will be established between function block D (the fourth leaf node) and the called class node Car2 (the second root node).

[0317] This represents the association between the class and the function block, indicating that this function block is part of the class; if there is a function call relationship between two leaf nodes (function blocks), then there will be a call edge between them, which means that one function block calls another function block; if the function called by a leaf node is not in the same class, then this leaf node will be marked as a call node, and the name of the call node is the function name stored in step S112. This node represents the function block's call to an external function.

[0318] Step S114: For function blocks without upper-level class nodes, a function call relationship is directly generated. At this time, if there is a call relationship between nodes SFCm and SFCn, there is a call edge between SFCm and SFCn, pointing from node SFCm to node SFCn; if there is a call relationship between node SFCm and the function within the class, node SFCm is pointed to class node SCL.

[0319] It should be understood that for those function blocks that are not in the class, a function call relationship is directly generated, that is, there is no class node in the upper layer, and the function block that has a call relationship with the function in the class is used as the fourth leaf node, and the class node called by the fourth leaf node is used as the second root node, indicating that these function blocks do not belong to any class, but are functions defined in the global scope, which means that these functions are visible and accessible in the entire program or code file, not just limited to a specific function or code block. These functions are usually defined at the top level of the program, that is, not defined in any other function, so they can be called in the entire program, which is different from local functions. Local functions are usually defined inside another function and can only be accessed inside the function that contains it. Global functions are usually used to perform tasks that need to be shared throughout the program, while local functions are used to implement more specific functions and are usually used in a specific context.

[0320] Optionally, if function block SFCm calls function block SFCn, then there will be a call edge in the generated graph, pointing from node SFCm to node SFCn, which means that function block SFCm calls function block SFCn in its internal code; if function block SFCm calls a function defined in a class node, then a call edge will be established, pointing from node SFCm to the corresponding class node SCL, which means that function block SFCm calls a function in the same class in its internal code. This section describes how to handle function blocks without upper-level class nodes and the calling relationships between them.

[0321] Step S115: Combine the nodes of step S114 and step S115 to generate a complete function call graph FCG. If there is a node in the node call of the function call graph FCG that is not in the stored function name, it is recorded as a call node. The principle is similar to step S902.

[0322] It should be understood that if the called function is not in the open source code file, the called function is saved as a call node based on the structured cache. For example, there are two source code files, file1.py and file2.py, and two functions, function_A and function_B. function_A is in file1.py, and function_B is in file2.py. Function_A calls function_B. Then, in the function call graph of file1.py, function_A is represented as a node. Since function_A calls function_B, this calling relationship needs to be represented. However, function_B is not in the same source code file file1.py, so function_B needs to be saved as a special call node instead of being regarded as a function block. Therefore, an edge is established from function_A to this special call node, indicating that function_A calls a certain function.

[0323] Optionally, step S115 first integrates the known function call relationships in the previous step S114, which includes the stored function names and the call relationships between them; after checking the known call relationships, step S115 starts to process the function names not stored in the previous steps, which may be functions defined in the global scope, or functions in other file sources; for the nodes of the unstored function names, these nodes will be recorded as "call nodes" to ensure that the system can track all function calls, including those unstored function names. Finally, S115 will generate a complete function call graph, which includes the stored function names, call nodes and the call relationships between them. Step S902 involves the identification of call nodes.

[0324] Optionally, for the open source software inter-process analysis process, a function call graph FCG is generated between the entire open source software. If the function call graph FCG of each file does not have a call node generated in step S115, it is considered that all source code files exist directly and independently and there is no call relationship.

[0325] It should be understood that step S115 is used to process function call relationships existing in open source software, which include calling known functions stored in previous steps, and calling unknown functions and functions that did not exist in storage in previous steps. If the call node generated in step S115 does not exist in the function call graph FCG of a source code file, it means that the function call relationship within this source code file will not be associated with the functions of other source code files. Therefore, it can be understood that the functions within this source code file will not call each other with the functions of other files. In analysis, this source code file is regarded as existing independently, and there is no cross-file function call between its internal functions.

[0326] Optionally, if there is a call node in the function call graph FCG, you need to find the called file through the following steps:

[0327] Step 1: Generate import nodes in the source code file where the call node exists, including general import nodes and hidden import nodes;

[0328] General type: import requests, which conform to the conventional import logic and are also the mainstream import logic in the industry;

[0329] Hiddenness: __import__("requests") does not conform to conventional import logic and is a complex way to implement simple logic.

[0330] It should be understood that Step 1 involves importing nodes, determining the target source code file that has a call node from other open source code files in the same open source software, and generating an import node for the target source code file; an import node is a part of the source code used to introduce external modules or libraries. In Python and other programming languages, importing modules or libraries is a common behavior that allows programmers to use code that has been written elsewhere for use in their own programs. Import nodes can include two types: general imports and hidden imports.

[0331] Optionally, "general import nodes" follow conventional, industry-wide accepted import logic, such as "importrequests", which is the most common import method, where "requests" is the name of a Python library or module that can be directly accessed and used in the code. General import nodes follow Python's standard import method and are usually automatically handled by the Python interpreter; "hidden import nodes" use unconventional or uncommon import logic, such as "__import__("requests")", which is different from general imports because it imports modules in a more flexible way at runtime. This import method is relatively uncommon and is usually used for special needs, such as dynamically loading modules, or loading different modules based on user input. Due to its non-standard nature, this import method is often called "hidden". Generally, hidden import nodes require more analysis because their behavior is more complex than general import nodes. When import nodes are identified, further analysis can determine which file calls these import nodes and whether these import nodes are associated with function call nodes.

[0332] Step 2: Find all import nodes, generate an abstract syntax tree, and traverse the edges between import nodes and call nodes in the tree. If the edge exists, the call file of the call node is the file corresponding to the import node, and the node is converted into a function node and deleted from the call node generated in step S115.

[0333] It should be understood that when an abstract syntax tree is generated based on an import node, if there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted.

[0334] It should be understood that Step 2 searches for all import nodes and generates an abstract syntax tree, that is, the import nodes in all source code files will be found, and then, for these import nodes, an abstract syntax tree will be generated. Next, the program will traverse the abstract syntax tree to find the edges between the import nodes and the call nodes. If an edge connecting the import node and the call node is found in the abstract syntax tree, then the call file of this call node is the file corresponding to the import node; in addition, this call node will be re-marked as a function node because it has been associated with a specific source code file. In this process, this call node will also be deleted from the call node list generated in step S115.

[0335] For example, Fig.12The flowchart of the process of converting a call node into a function node in a vulnerability identification method provided by the present application is shown as follows. Consider a function call graph including function call nodes and import nodes, which may involve a function call of the import module requests. In the initial state: (1) Node A is a function call node, which calls the module "requests"; (2) Node B is an import node, which imports the requests module. Deletion process: (1) Identify node B as an import node, and determine that the call file of node A is the file corresponding to node B; (2) Change the type of node A from a function call node to a function node; (3) Delete node A from the function call graph because it is now considered a function node. This is just a simplified example. The actual process may be more complicated, depending on the specific implementation of the graph database and data structure. In the graph database, these operations are usually handled programmatically to ensure data consistency and accuracy.

[0336] Step 3: Generate the function call graph FCG of the entire open source software by linking all the call nodes. The function call graph FCG is not a graph with a call path between any two points. There are some independent function call graph FCG subgraphs, which have no nodes connected to other subgraphs.

[0337] It should be understood that Step 3 generates the function call graph FCG of the entire open source software by connecting all the call nodes. This function call graph FCG can be regarded as composed of multiple subgraphs, each of which represents the function call relationship in a file. However, not all subgraphs have direct call paths between nodes, and some subgraphs may be relatively independent. This step helps to clarify the association relationship of function calls, connect function nodes and file nodes, and form a function call graph in the entire open source software.

[0338] The above steps are to better understand the function call relationship within the open source software, ensure that all call nodes can be correctly associated with the corresponding files, and help generate the function call graph FCG of the entire software for further analysis and vulnerability detection.

[0339] against Figure 2 Step S243 in the method: open source software dependency analysis. Fig.13 A schematic diagram of generating an open source dependency graph in a vulnerability identification method provided by the present application is shown for illustration:

[0340] If the FCG of an open source software does not have the calling node generated in step S115, then the open source software does not have a calling relationship, but only a called relationship.

[0341] It should be understood that in a function call graph of an open source software, the calling relationship represents the interaction between functions. If function A calls function B, then there is a calling edge from function A to function B, which can be understood as function A calling function B in some way, and there is a calling relationship between them; the called relationship represents a function being called by other functions. If function B is called by function A, then there is a called edge from function A to function B, which can be understood as function B being called in function A, and there is a called relationship between them.

[0342] Optionally, if the function call graph of an open source software does not contain the call node generated in step S115, but there are other function nodes, this means that the open source software contains a set of functions, some of which are called by other functions but do not call other functions in turn. It can be understood that this open source software mainly contains the called relationship of functions but does not have a significant active calling relationship between functions.

[0343] In practice, this situation may occur in library functions or general functions, which are referenced by other parts of the code, but they do not call other functions themselves. It should be noted that this does not mean that the open source software has no calling relationship, but that the calling relationship is relatively small, and the called relationship may be more significant.

[0344] For open source software with call nodes, extract and search all import nodes in the source code file to generate an abstract syntax tree, and traverse the edges between the import nodes and the call nodes in the tree. If the edge exists, the calling open source software of the call node is the open source software corresponding to the import node, and the node is converted into a function node, deleted from the call nodes generated in step S115, and the call nodes in step S115 are traversed in a loop until the call node is empty and the traversal ends.

[0345] It should be understood that merging the generated function call graphs to generate software dependency entities may include: based on the function call graphs generated by each open source software, for the open source software with call nodes, extracting import nodes from other open source software in the open source software dependency list; generating an abstract syntax tree based on the extracted import nodes, if there is an edge connecting the import node and the call node, recording the call node as a function node, and deleting the call node saved based on the structured cache; until all call nodes in the open source software with call nodes are processed.

[0346] It should be understood that the function call graph of open source software includes call nodes, which represent the mutual call relationship between functions. The source code files in the open source software are scanned to find the import nodes therein. The import nodes are usually used to import other modules or libraries for use, and an abstract syntax tree is generated. The abstract syntax tree is traversed, and the edges between the import nodes and the call nodes are paid special attention to. The edges here represent that the import nodes refer to or use the functions of other modules or libraries. If an edge from the import node to the call node is found during the traversal process, it means that the call node actually calls the imported module or library. This process links the call node with the open source software they actually call. At this point, the call node is no longer a pure call node because it is now associated with a specific open source software. Therefore, the type of the node is changed to a function node to reflect its actual behavior, and at the same time, the node is deleted from the original call node set. Repeat the above process until all call nodes have been processed. This cyclic process ensures that all functions with call nodes can be correctly associated with the open source software they call.

[0347] like Fig.13 As shown, all source code files in the software product are traversed to find all import nodes. After determining the open source software corresponding to the import node, an abstract syntax tree is generated for the source code file. The import node is regarded as the root node, and all leaf nodes are traversed to obtain the leaf nodes. The leaf nodes are mapped to the function nodes in the open source software to obtain a complete open source dependency graph including the software product and the open source software. The dark black nodes are the leaf nodes with the calling relationship, and the black dotted lines are the calling edges.

[0348] It should be understood that this process describes how to create a complete open source dependency graph in a software product, mapping the source code files of the software product with the function nodes of the open source software to present the relationship between the software product and its open source dependencies.

[0349] First, you need to traverse all source code files in the software product, including directories and subdirectories of the source code files. In each source code file, you need to find all import nodes. Import nodes are usually used to introduce external modules, libraries, or software packages so that their functions can be used in the current file.

[0350] Secondly, for each import node, it is necessary to determine which open source software the external module, library or software package it introduces is. This can usually be determined by the module name or library name specified in the import node. Once the open source software corresponding to the import node is determined, it is necessary to generate an abstract syntax tree for the current source code file. This abstract syntax tree can reflect the structure and hierarchy of the current source code file. The import node is regarded as the root node, and the abstract syntax tree is traversed from this root node to obtain the leaf nodes in the tree. These leaf nodes represent specific code segments in the source code file.

[0351] Finally, for each leaf node, it needs to be mapped to the function node in the open source software. This mapping indicates which open source software function is used by the source code file of the software product. As the source code files are traversed and the leaf nodes are mapped, a complete open source dependency graph containing software products and open source software will be constructed. This graph shows the relationship between the software product and the open source software it depends on, including which functions are called by the software product's code.

[0352] The above process helps to understand the open source dependencies of software products and clarify which open source software functions are used by software products, which is very important for software product maintenance, vulnerability detection and security assessment.

[0353] against Figure 2 Step S244 in the method: vectorized storage of the dependency graph.

[0354] It should be understood that storing the open source dependency graph in a graph database (such as Neo4j) is to effectively manage and query the relationships and dependencies between open source software.

[0355] First, establish a connection. You need to establish a connection with the graph database. Usually you need to provide connection information such as the database host address, port number, user name and password. Once the connection is established, you can communicate with the database.

[0356] Second, create a graph. In Neo4j, a database can contain multiple graphs. You can create a new graph and use it to store open source dependencies. This can be performed through database management tools or database clients, or programmatically through the graph database's API.

[0357] Third, create nodes. Before starting to store the open source dependency graph, you must first create nodes. Nodes represent open source software, source code files, functions, classes, etc. Each node should include necessary attributes, such as name, version, type, etc. These attributes will become the basis for subsequent queries and associations.

[0358] Fourth, add nodes to the graph. Once the nodes are created, they can be added to the graph. This can be done by executing the corresponding commands of the graph database or programming using the API. Each node should be associated with the graph.

[0359] Fifth, create relationships and the relationships between associated nodes. Dependencies are the core of the open source dependency graph. Different types of relationships need to be defined to represent dependencies between open source software, reference relationships between source code files, calling relationships between functions, etc. These relationships usually have attributes to describe the details of the relationship; for each relationship, it is necessary to clearly identify the two nodes it connects. For example, a relationship may connect an open source software node and a source code file node, indicating that the software depends on this file.

[0360] Sixth, add relationships to the graph. Similar to nodes, relationships need to be added to the graph, which will establish the relationship between nodes.

[0361] Seventh, perform queries and analysis. Once the graph data is stored, you can use query languages ​​(such as Cypher for Neo4j) to perform various queries and analyses. Cypher is a graph database query language specifically used for graph database systems such as Neo4j. Cypher is a declarative query language designed for retrieving, inserting, updating, and deleting data in graph databases. These queries can help understand the relationships between open source software and find specific dependencies.

[0362] Eighth, maintenance and updating. The dependencies of open source software may change over time. Therefore, the graph data storage needs to be regularly maintained and updated to ensure that it is consistent with the actual situation.

[0363] against Figure 2 Step S250 in the method: security vulnerability knowledge graph. This can be specifically achieved by:

[0364] Step S251: based on the open source dependency graph generated in step S244.

[0365] Step S252: Based on the vulnerability knowledge generated in step S234.

[0366] Step S253: Based on the open source dependency graph of step S251 and the vulnerability knowledge of step S252, a security vulnerability knowledge graph of the software product to be analyzed is generated. The security vulnerability knowledge graph of the software product to be analyzed includes vulnerability nodes, software nodes of the software product to be analyzed, calling relationships between the software product to be analyzed and open source software, and dependencies between open source software. The calling relationships include calling relationships between functions in the software code.

[0367] It should be understood that the vulnerability node represents a known security vulnerability, which may generally include the type, name, description, severity rating and recommended repair measures of the vulnerability. Each vulnerability node is used to represent a specific vulnerability. The software node of the software product to be analyzed represents the target software product to be analyzed, which may generally include source code files, classes, functions, statements and other elements in the software product. These nodes help to determine which vulnerabilities exist in the target software. The calling relationship between the software product to be analyzed and the open source software indicates the existence of a calling relationship between the target software product and the functions in the open source software, that is, how one function calls another function, and there are edges between them. These edges in the dependency relationship between open source software represent the dependency relationship between open source software, including which open source software is used to build the target software, and their versions.

[0368] Optionally, the open source dependency graph is defined as a software dependency entity, and the generated function call graphs are merged to generate a software dependency entity, the software dependency entity includes the dependency relationship between the open source software, and the dependency relationship between the open source software includes the dependency relationship of each function call graph.

[0369] It should be understood that the various components in the open source dependency graph (dependencies between open source software, various function call graphs, etc.) are regarded as software dependency entities. The software dependency entity represents the entire software product to be analyzed and the open source software it depends on, as well as the dependencies between them, which may include the dependencies between the software product to be analyzed and the open source software, as well as the dependencies between the open source software. These dependencies may include dependencies between various function call graphs.

[0370] Vulnerability knowledge can be defined as vulnerability entities. The two major entities contain multiple sub-entity structures, and different entities are connected through different association relationships.

[0371] Security vulnerability knowledge graph = software dependency entity + vulnerability entity

[0372] Software dependent entities include:

[0373] Nodes: class nodes, function nodes, statement nodes in source code files, call nodes between source code files, dependency nodes between open source software, and dependency nodes between software products and open source software;

[0374] Edges: CFG edge, FCG edge, call edge, dependency edge.

[0375] Vulnerable entities include:

[0376] Nodes: affected file nodes, affected class nodes, affected function nodes, affected statement nodes;

[0377] Edge: has_*, whether there is an execution order relationship between nodes.

[0378] In the security vulnerability knowledge graph, G is used to represent the knowledge graph, V is used to represent the node set, and E is used to represent the edge set. The formal description of the knowledge graph is as follows:

[0379] G = {V, E}

[0380] Formal description of nodes:

[0381] V={v1,v2,…,v i ,…,v m}|1≤i≤m|

[0382] Formal description of edges:

[0383] E={E + ,E -}

[0384] E + ={e 1,1 ,e 1,2 ,…,e 1,m ,e 2,1 ,e 2,2 ,…,e 2,m ,…,e m,1 ,e m,2 ,…,e m,m}

[0385] E - ={-e 1,1 ,-e 1,2 ,…,-e 1,m ,-e 2,1 ,-e 2,2 ,…,-e 2,m ,…,-e m,1 ,-e m,2 ,…,-e m,m}

[0386] E is an edge set, including forward adjacent edges and reverse adjacent edges. E+ is the forward adjacent edge set, ei,j represents the edge from node vi to node vj; E- is the reverse adjacent edge set, -ei,j represents the edge from node vj to node vi.

[0387] The security vulnerability knowledge graph is a directed graph. To detect whether a software product has open source vulnerabilities introduced by open source software, it is necessary to find out whether there is reachability between the software product node and the vulnerability node.

[0388] It should be understood that a directed graph is a graph in which the relationship between nodes is directional, and the nodes represent various software products, such as operating systems, applications, etc. Vulnerability nodes represent known security vulnerabilities and may include: vulnerability information, open source software name, affected version number, affected files, affected classes, affected functions, affected statements and other information.

[0389] Optionally, when detecting whether a software product has an open source vulnerability introduced by open source software, it is necessary to find out whether there is reachability between the software product node and the vulnerability node, that is, whether there is a directed path from the software product node to the vulnerability node. If such a path exists, it means that the software product is threatened by the vulnerability because the open source software containing the vulnerability is used in the software product. For example, reachability: if there is an edge from node A to node B, it is considered that A is reachable from B.

[0390] Fig.14 This is a flow chart of identifying the call chain path of an open source vulnerability in a vulnerability identification method provided by this application, such as Fig.14 As shown, a graph search algorithm is used to search the security vulnerability knowledge graph, with the vulnerability node as the starting node, and the reachability of each vertex in the node set is determined by recursive marking.

[0391] It should be understood that the graph search algorithm is used to find a specific node in the graph or to find other related nodes from a starting node. In the case of the application, the goal is to find which other nodes can be reached from a specific vulnerability node. It can be understood that the vulnerability node is the starting node of the search, representing a specific vulnerability. The search will start from this node and find other nodes related to it.

[0392] It should be understood that the recursive marking bit is a marking method used to mark the reachability of nodes during the search process. When a node is marked as reachable, it means that the node can be reached from the vulnerable node. During the search process, the reachability of the node set needs to be determined; the search algorithm will traverse each node and determine whether these nodes can be reached from the vulnerable node through the recursive marking bit. If they can be reached, these nodes will be marked as reachable. This helps to identify other nodes related to a specific vulnerability, such as affected software products or vulnerable code segments, so as to better understand the scope of the vulnerability and take corresponding security measures.

[0393] Step S141: Mark vulnerability nodes and software nodes in the security vulnerability knowledge graph.

[0394] It should be understood that in the security vulnerability knowledge graph, nodes are the basic building blocks of the graph, each node represents a specific entity or object, and "vulnerability nodes" and "software nodes" are two different types of nodes; vulnerability nodes represent different security vulnerabilities, and each vulnerability node contains information related to a specific vulnerability, such as vulnerability information, open source software name, affected version number, affected files, affected classes, affected functions, affected statements, severity, disclosure date, etc.; software nodes represent class nodes, function nodes, statement nodes within source code files, call nodes between source code files, dependency nodes between open source software, and dependency nodes between software products and open source software. ,Software nodes are used to represent the structure and relationships within the entire software system, as well as the dependencies between the software and external entities (such as software products and open source software). The structure within the source code file can include class nodes, function nodes, statement nodes, and the relationships and call relationships between them. The dependencies between source code files represent the dependencies between source code files, for example, one file may introduce the functions or libraries of another file. Open source software dependencies represent the open source software used by the software product, and the dependencies between them, including version information and other related data. Software product dependencies represent the dependencies between software products and other software products, such as operating systems, libraries, services, etc.

[0395] Step S142: Use a graph search algorithm to find a reachable path from the vulnerability node to the software node in the security vulnerability knowledge graph.

[0396] It should be understood that the graph search algorithm can be a breadth-first search (BFS), a depth-first search (DFS) or a more complex algorithm, such as the Dijkstra algorithm or the A* search algorithm, which are used to find a path from a vulnerability node to a software node in a graph; a reachable path means that there are one or more associated nodes (usually vulnerability nodes or software nodes) connected to each other from the vulnerability node to the software node, forming a path. The path may pass through multiple nodes, but it represents how the vulnerability spreads to the software product.

[0397] Optionally, weights can be introduced for edges and nodes in the graph to represent the importance of different relationships, which helps identify paths with higher potential risks; different data sources such as vulnerability databases, code review tools, and dependency analysis tools can be combined to obtain more context for more precise graph searches.

[0398] The reachable path of step S142 is the open source vulnerability call chain path.

[0399] It should be understood that vulnerabilities are not isolated, they exist in the software system and propagate through the calling relationship between functions, classes, files or modules. The vulnerability call chain path describes how the vulnerability propagates from one part to another; the reachable path usually refers to the critical path, which is the main steps that a vulnerability goes through when it propagates, including the call sequence of specific functions, methods, modules, etc., which ultimately leads to the manifestation of the vulnerability. By understanding the reachable path, the scope of the vulnerability's possible impact can be determined. This is crucial for assessing the severity of the vulnerability and determining the importance of fixing it. The reachable path or vulnerability call chain path is a key part of security analysis. They help understand how vulnerabilities propagate, locate vulnerabilities, determine repair strategies, and improve the security of software systems.

[0400] against Figure 2 Step S260 in the method shown: open source vulnerability warning and risk assessment. It may specifically include:

[0401] Based on the scanning results of the vulnerability scanning core components, open source vulnerability alerts and risk assessments are generated, indicating which open source software the software product relies on has open source vulnerabilities and the degree of harm of the open source vulnerabilities; and the information is pushed to internal developers for rectification and repair.

[0402] It should be understood that the vulnerability scanning core component is the core engine of the open source vulnerability analysis system, which is responsible for analyzing the source code, dependencies and configurations of software products to detect potential vulnerabilities and identify open source vulnerability call chain paths, including vulnerability knowledge extraction, parallel generation of open source dependency graphs, and generation of security vulnerability knowledge graphs.

[0403] Optionally, once the system scan is completed, it will generate warning information about vulnerabilities in the open source software. This information usually includes the type of vulnerability, name, affected version, description and recommended repair measures. Vulnerabilities are not equal. Some vulnerabilities may pose a greater threat to the security and stability of the system. Risk assessment considers the severity of the vulnerability and usually uses a standard assessment system, such as CVSS, to determine the degree of harm caused by the vulnerability.

[0404] Based on the above content, the ultimate goal of the open source vulnerability analysis system is to ensure that potential vulnerabilities are fixed. Therefore, the system pushes the generated alerts and risk assessment information to internal developers or teams so that they can take appropriate solutions.

[0405] Optionally, some open source vulnerability analysis systems have automated repair capabilities and can provide repair suggestions or even automatically repair some common vulnerabilities, which can significantly reduce the complexity and workload of repairs; continuous integration / continuous delivery (CI / CD), vulnerability scanning can be integrated into the CI / CD process to achieve automated vulnerability detection and repair, which helps ensure the security of new code and discover and resolve vulnerabilities before release; open source vulnerability analysis systems can be integrated with vulnerability management systems to track the status and repair progress of vulnerabilities, which is very important for team collaboration and tracking the progress of vulnerability repairs.

[0406] In summary, a security vulnerability knowledge graph was constructed, all open source software call chains were traversed based on a graph search algorithm, and a graph reachability algorithm was used to simulate the propagation path of open source software vulnerabilities between open source software dependency levels. In addition, the propagation path of open source vulnerabilities within open source software was located based on a named entity recognition model and a path sorting algorithm. Among them, inter-procedural analysis, intra-procedural analysis, and dependency analysis were performed on the open source software source code based on a parallel algorithm, and the calling relationship between software products and open source software was constructed at the function granularity, so as to accurately determine whether open source vulnerabilities will affect software products and reduce the interference caused by invalid components and vulnerability information.

[0407] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the information identification method.

[0408] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by the computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a high-density digital video disc (DVD)), or a semiconductor medium (for example, a solid-state hard disk), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the information identification method, or instruct the computing device to execute the information identification method.

[0409] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes a plurality of computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0410] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent repairs or replacements within the technical scope disclosed by the present invention, and these repairs or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. An open source vulnerability analysis method, characterized in that: include: Generate a knowledge graph of security vulnerabilities of the software product to be analyzed; The security vulnerability knowledge graph includes vulnerability nodes, software nodes of the software product to be analyzed, calling relationships between the software product to be analyzed and open source software, and dependency relationships between open source software; the calling relationships include calling relationships between functions in the software code; Scanning a reachable path from a vulnerability node to the software node in a security vulnerability knowledge graph based on a graph search algorithm; Generate vulnerability analysis results based on the scan results.

2. The method according to claim 1, characterized in that The step of generating a security vulnerability knowledge graph of the software product to be analyzed includes: Obtaining an open source software dependency list for the software product to be analyzed; the open source software dependency list includes a plurality of open source software that the software product to be analyzed depends on; Generate a function call graph for each open source software in the open source software dependency list; The generated function call graphs are merged to generate a software dependency entity; the software dependency entity includes the dependency relationship between the open source software, and the dependency relationship between the open source software includes the dependency relationship of each function call graph.

3. The method according to claim 2, characterized in that The generating of a function call graph for each open source software in the open source software dependency list includes: For each open source software in the open source software dependency list, decompose each open source code file in the open source software; For each open source code file in the open source software, decompose each class in the open source code file; Decompose each function in the class according to the code snippet of each class in the open source code file; For each code snippet of the function, decompose the statements contained in the function; For each statement in the same function, a control flow graph is generated with the statement as a node; if the statement has the behavior of calling other functions, the control flow graph includes a call node saved based on the structured cache, and there is an edge between the statement node that calls other functions and the call node; Generate a function call graph within the same class based on the control flow graph; Merge the function call graphs of all classes in the same open source code file to generate a function call graph of the open source code file; The function call graphs of each open source code file in the same open source software are merged to generate a respective function call graph of each open source software in the open source software dependency list.

4. The method according to claim 3, characterized in that For each statement in the same function, a control flow graph is generated with the statement as a node, including: Divide the function into statements, with each statement as a statement node; If the statement contains a behavior of calling other functions, determining the calling node, and saving the calling node based on the structured cache; A control flow graph is generated based on statement nodes and call nodes; the edges between nodes in the control flow graph represent the execution order of the nodes or the call relationship between the nodes.

5. The method according to claim 4, characterized in that The determining of the calling node includes: if there is a hidden node, traversing upward with the hidden node as a leaf node, and taking the found root node as the calling node; and / or, Different data dependency analyses are performed on calls with different parameters to determine the call nodes.

6. The method according to claim 3, characterized in that generating a function call graph within the same class based on the control flow graph; Merge the function call graphs of all classes in the same open source code file to generate the function call graph of the open source code file; including: Get all function blocks in the same open source code file; Based on the control flow graph, an edge is generated between the first root node and the first leaf node, an edge is generated between the second leaf node and the third leaf node, and an edge is generated between the fourth leaf node and the second root node; Among them, the function block with a class node in the upper layer is the first leaf node, and the class node with a function block below is the first root node; the function block that calls other functions is the second leaf node, and the function block called by other functions is the third leaf node; the function block that has no class node in the upper layer and has a calling relationship with the function in the class is the fourth leaf node, and the class node called by the fourth leaf node is the second root node; If the called function is not in the open source code file, the called function is saved as a calling node based on the structured cache.

7. The method according to claim 6, characterized in that The step of merging the function call graphs of each open source code file in the same open source software to generate a function call graph for each open source software in the open source software dependency list includes: For the calling node, determining a target source code file of the calling node from other open source code files in the same open source software; Generate an import node for the target source code file; An abstract syntax tree is generated based on the import node. If there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted.

8. The method according to claim 7, characterized in that The step of merging the generated function call graphs to generate a software dependency entity includes: Based on the function call graph generated by each open source software, for the open source software with call nodes, extracting import nodes from other open source software in the dependency list of the open source software; An abstract syntax tree is generated based on the extracted import node. If there is an edge connecting the import node and the call node, the call node is recorded as a function node, and the call node saved based on the structured cache is deleted; until all the call nodes in the open source software with the call node are processed.

9. The method according to any one of claims 2 to 8, characterized in that: The generating of the security vulnerability knowledge graph of the software product to be analyzed also includes: Associating a vulnerability node where a software vulnerability exists with the software dependency entity; Based on the execution order involved in the vulnerability nodes, the calling relationship of each node in the security vulnerability knowledge graph is integrated.

10. The method according to claim 9, characterized in that Before associating the vulnerability node with the software vulnerability on the software dependency entity, the method further includes: For each vulnerable open source software, obtain an open source vulnerability source code fragment of the vulnerable open source software; Analyze the open source vulnerability source code fragment to generate a vulnerability knowledge graph of the vulnerable open source software; the vulnerability knowledge graph includes name information, version information, vulnerability statements and vulnerability propagation path of the vulnerable open source software, and the vulnerability propagation path includes a propagation path from a root node to a node of the vulnerability statement; Integrate the vulnerability knowledge graphs of multiple vulnerable open source software.

11. The method according to claim 10, characterized in that The step of obtaining the open source vulnerability source code fragment of the vulnerable open source software includes: Based on the open source vulnerability library, vulnerability knowledge is extracted from the unstructured text in the open source vulnerability web pages; Extract open source vulnerability source code fragments of vulnerable open source software from the vulnerability knowledge.

12. The method according to claim 11, characterized in that The analyzing the open source vulnerability source code fragment includes: Slicing the open source vulnerability source code fragment, parsing the source code fragment into a syntax node of an abstract syntax tree, adding node labels to the nodes of the abstract syntax tree, mapping the source code fragment to the abstract syntax tree, and adding instance nodes to the abstract syntax tree; The path of the leaf nodes of the abstract syntax tree is traversed to generate a propagation path from the root node to the leaf nodes; the leaf nodes are label nodes marked with open source vulnerabilities.

13. The method according to claim 11 or 12, characterized in that: After extracting vulnerability knowledge from unstructured text in the open source vulnerability web page based on the open source vulnerability library, the method further includes: Extracting an open source vulnerability description statement of the vulnerable open source software from the vulnerability knowledge; Performing text parsing on the open source vulnerability description statement to obtain vulnerability named entity information; the vulnerability named entity information includes the name information, version information, vulnerability statement of the vulnerable open source software and at least one of the following: vulnerability file information, vulnerability class, and vulnerability function; Wherein, in the process of generating the vulnerability knowledge graph of the vulnerable open source software, the source code information extracted by analyzing the open source vulnerability source code fragment is integrated with the vulnerability named entity information.

14. The method according to claim 13, characterized in that The text parsing of the open source vulnerability description statement to obtain vulnerability named entity information includes: Performing word segmentation processing on the open source vulnerability description sentence to obtain a sequence after word segmentation; The sequence after word segmentation is processed through a bidirectional encoder representation network based on Transformer to obtain word encoding, sentence encoding and position encoding; The word code, the sentence code and the position code are summed and combined; The merged vector encoding is processed through a bidirectional long short-term memory model and conditional random field to output vulnerability named entity information.

15. An open source vulnerability analysis device, characterized in that: Comprising means for performing the steps of the method as claimed in any one of claims 1 to 14.

16. An open source vulnerability analysis device, characterized in that: include: one or more processors and one or more memories; The one or more memories are coupled to one or more processors, and the one or more memories are used to store computer executable programs. When the one or more processors execute the computer executable programs, the electronic device executes the method as described in any one of claims 1-14.

17. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 14.

Citation Information

Cited By

  • Software vulnerability detection method and device, equipment, medium and program product

    CN120995472A

  • Method and device for automatically generating vulnerability analysis report of database component and medium

    CN121234378A

  • Security intelligent monitoring and risk assessment method for open source software supply chain

    CN121580391A