Component security detection method and device, electronic equipment and readable storage medium
By static analysis and code tree comparison of components in the software supply chain, the problem of undisclosed security vulnerabilities and attack codes in the prior art is solved, and the accuracy of component security detection is improved.
Patent Information
- Application Number
- CN202510226794.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When detecting the security of components in the software supply chain, the prior art cannot identify components that have not been disclosed as having security vulnerabilities, and the attack code inserted into the components cannot be detected, resulting in low detection accuracy.
By statically analyzing the source code of the component to be detected, a code tree is built, and tree similarity comparison is performed with the standard component code tree to determine the security of the component. The method includes parsing the source code into an abstract syntax tree, deleting the target noise type node, and calculating the hierarchical feature density and similarity level to evaluate the security of the component.
Improves the accuracy of component security detection, can identify undisclosed security vulnerabilities and components mixed with attack code, and enhances the overall security of the software supply chain.
Smart Images

Figure CN120162780A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software security technology, and in particular, to a method, device, electronic device, and readable storage medium for detecting component security. Background Art
[0002] In today's increasingly complex software development ecosystem, software supply chain security has become a key issue that cannot be ignored. The software supply chain covers every component that interacts with software between the development stage and the release process of a software application. Components are the basic units in the software supply chain and are used to build software applications. These components may include open-source libraries, third-party frameworks, custom modules, etc.
[0003] The security of components is crucial to the overall security and stability of the software supply chain and the built software applications. If there are vulnerabilities or security issues in the components, it may pose a threat to the entire software supply chain and the built software applications. Therefore, in software supply chain management and software application security management, it is necessary to detect the security of components.
[0004] Currently, when detecting the security of components, it is mainly to detect the file list included in the software application. The file list contains the specifications of each component in the software application, and the specifications of the components record whether the component has been publicly confirmed to have security vulnerabilities. That is to say, in the prior art, it is only possible to detect whether the software application contains components that have been publicly confirmed to have security vulnerabilities by detecting the specifications of each component. However, some components (such as component A) actually have security vulnerabilities but have not been made public, that is, the specification does not record that component A has been publicly confirmed to have security vulnerabilities. At this time, through the detection method of the prior art, it is impossible to identify that component A has security vulnerabilities, resulting in low accuracy of component security detection.
[0005] Furthermore, since the detection method in the prior art is to detect the specifications of each component and cannot detect the code of the components, and some normal components are inserted with attack code, it is impossible to detect the components mixed with attack code through the detection method of the prior art, which also results in low accuracy of component security detection. Summary of the Invention
[0006] In view of this, the purpose of this application is to provide a method, device, electronic device, and readable storage medium for detecting component security to improve the accuracy of component security detection.
[0007] In a first aspect, an embodiment of this application provides a method for detecting component security, and the method includes:
[0008] Perform static analysis on the source code of the component to be detected, and construct a code tree of the component to be detected for the component to be detected;
[0009] Compare the code tree of the component to be detected with the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree; wherein, the standard component code tree is constructed by performing static analysis on the source code of the standard component; the standard component is the standard component of the component to be detected;
[0010] Determine the security of the component to be detected based on the similarity.
[0011] Combined with the first aspect, the embodiment of the present application provides a first possible implementation manner of the first aspect, wherein, performing static analysis on the source code of the component to be detected and constructing a code tree of the component to be detected for the component to be detected includes:
[0012] Input the source code of the component to be detected into the ast module, and parse the source code of the component to be detected into an abstract syntax tree through the ast module; wherein, the abstract syntax tree contains multiple nodes, and each node corresponds to its own node type;
[0013] Delete the nodes with the target noise type in the abstract syntax tree from the abstract syntax tree to obtain the code tree of the component to be detected for the component to be detected; wherein, the target noise type includes: expression node, string node, number node, name constant node, placeholder node.
[0014] Combined with the first aspect, the embodiment of the present application provides a second possible implementation manner of the first aspect, wherein, comparing the code tree of the component to be detected with the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree includes:
[0015] For each level in the code tree of the component to be detected, calculate the level feature density of this level;
[0016] Calculate the similarity degree of the same level in the code tree of the component to be detected and the standard component code tree respectively to obtain the similarity degree of each level in the code tree of the component to be detected and the standard component code tree;
[0017] Calculate the similarity between the code tree of the component to be detected and the standard component code tree according to the similarity degree corresponding to each level in the code tree of the component to be detected and the standard component code tree, and the level feature density corresponding to each level.
[0018] Combined with the second possible implementation manner of the first aspect, the embodiments of the present application provide a third possible implementation manner of the first aspect, wherein calculating the hierarchical feature density of each level in the code tree of the component to be detected includes:
[0019] For each node in the i-th level of the code tree of the component to be detected, determine the parent-child relationship feature between the node and its parent node according to the node type feature of the node and the node type feature of the parent node of the node, so as to obtain the parent-child relationship feature between each node in the i-th level and its respective parent node; wherein, the value range of i is from 1 to L, and there are L layers in the code tree of the component to be detected except the root node; when i takes 1, each node in the i-th level is a child node of the root node;
[0020] Aggregate the parent-child relationship features between each node in the i-th level and its respective parent node, and perform a hash calculation on the aggregation result to obtain the fingerprint of the i-th level;
[0021] Calculate the node type Shannon entropy of the node according to the probability of the node type of each child node of the node among all the child nodes of the node;
[0022] Calculate the node type Shannon entropy of each ancestor node of the node respectively, and calculate the depth-aware ancestor entropy of the node according to the node type Shannon entropy of each ancestor node of the node;
[0023] Calculate the adjacent hierarchical structure relationship smooth entropy corresponding to the node according to the fingerprint of the i-th level where the node is located and the fingerprint of the (i - 1)-th level where the parent node of the node is located;
[0024] Calculate the fusion comprehensive entropy of the node based on the node type Shannon entropy of the node, the depth-aware ancestor entropy of the node, the adjacent hierarchical structure relationship smooth entropy corresponding to the node, and a preset weight coefficient;
[0025] For each node in the i-th level of the code tree of the component to be detected, calculate the first ratio of the fusion comprehensive entropy of the node to the sum value of the fusion comprehensive entropy of each node in the i-th level, and calculate the sampling probability of the node according to the first ratio and a preset first attenuation coefficient;
[0026] According to the sampling probability of each node in the i-th level of the code tree of the component to be detected, select some nodes from each node in the i-th level as target nodes;
[0027] Calculate the initial hierarchical feature density of the i-th level based on the fingerprint of the i-th level, the absolute value of the set of node type features of all target nodes in the i-th level, and the maximum value among the absolute values corresponding to each level of the code tree of the component to be detected; wherein, the absolute value corresponding to a level is the absolute value of the set of node type features of all target nodes within that level.
[0028] Normalize the initial hierarchical feature density of the i-th level to obtain the hierarchical feature density of the i-th level.
[0029] Combined with the third possible implementation manner of the first aspect, the embodiments of the present application provide a fourth possible implementation manner of the first aspect, wherein, the normalizing the initial hierarchical feature density of the i-th level to obtain the hierarchical feature density of the i-th level includes:
[0030] Calculate the second ratio of the initial hierarchical feature density of the i-th level to the sum value of the initial hierarchical feature densities of each level in the code tree of the component to be detected.
[0031] Calculate the hierarchical feature density of the i-th level according to the second ratio and a preset second attenuation coefficient.
[0032] Combined with the third possible implementation manner of the first aspect, the embodiments of the present application provide a fifth possible implementation manner of the first aspect, wherein, the calculating the similarity degree of the same level in the code tree of the component to be detected and the code tree of the standard component respectively includes:
[0033] For each node in the i-th level of the code tree of the component to be detected, perform a preset number of hash calculations on the node type feature of this node to obtain the hash value of this node, and perform the preset number of hash calculations on the node type feature of each child node of this node to obtain the hash values of each child node, and calculate the sum value of the hash value of this node and the hash values of each child node.
[0034] Perform an exclusive OR calculation on the sum value of the hash value of this node and the hash values of each child node, and the fingerprint of the relationship feature between this node and its parent node, to obtain the matrix corresponding to this node; wherein, the matrix contains multiple element values.
[0035] Determine the minimum element value in the matrix corresponding to this node.
[0036] After obtaining the minimum element value in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, and use this minimum value as the first MinHash signature corresponding to the i-th level of the code tree of the component to be detected.
[0037] For the i-th level, calculate the similarity degree of the i-th level between the code tree of the component to be detected and the standard component code tree according to the first MinHash signature corresponding to the i-th level of the code tree of the component to be detected and the second MinHash signature corresponding to the i-th level of the standard component code tree.
[0038] Combined with the first aspect, the embodiment of the present application provides a sixth possible implementation manner of the first aspect. Before statically analyzing the source code of the component to be detected and constructing the code tree of the component to be detected, the method further includes:
[0039] Recursively process the component code file of the component to be detected, traverse the component code file, and obtain the file-level architecture of the component code file;
[0040] Parse the file-level architecture to obtain each sub-file in the file-level architecture;
[0041] For each sub-file, if the source code contained in the sub-file is bytecode, decompile the bytecode to obtain the source code of the sub-file;
[0042] Determine the source code of all the sub-files in the file-level architecture as the source code of the component to be detected.
[0043] In a second aspect, the embodiment of the present application further provides a component security detection device, and the device includes:
[0044] A construction module, configured to statically analyze the source code of the component to be detected and construct a code tree of the component to be detected;
[0045] A comparison module, configured to perform a tree similarity comparison between the code tree of the component to be detected and the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree; wherein, the standard component code tree is constructed by statically analyzing the source code of the standard component; the standard component is the standard component of the component to be detected;
[0046] A first determination module, configured to determine the security of the component to be detected based on the similarity.
[0047] Combined with the second aspect, the embodiment of the present application provides a first possible implementation manner of the second aspect. When the construction module is configured to statically analyze the source code of the component to be detected and construct the code tree of the component to be detected, it is specifically configured to:
[0048] Input the source code of the component to be detected into the ast module, and parse the source code of the component to be detected into an abstract syntax tree through the ast module; wherein, the abstract syntax tree contains multiple nodes, and each node corresponds to its own node type;
[0049] Delete the nodes with the target noise type in the abstract syntax tree from the abstract syntax tree to obtain the code tree of the component to be detected of the component to be detected; wherein, the target noise types include: expression nodes, string nodes, numeric nodes, name constant nodes, placeholder nodes.
[0050] Combined with the second aspect, the embodiment of the present application provides a second possible implementation manner of the second aspect. When the comparison module is used to compare the code tree of the component to be detected with the code tree of the standard component to obtain the similarity between the code tree of the component to be detected and the code tree of the standard component, it is specifically used for:
[0051] For each level in the code tree of the component to be detected, calculate the level feature density of this level;
[0052] Calculate the similarity degree of the same level in the code tree of the component to be detected and the code tree of the standard component respectively, so as to obtain the similarity degree of each level in the code tree of the component to be detected and the code tree of the standard component;
[0053] According to the similarity degree corresponding to each level in the code tree of the component to be detected and the code tree of the standard component, and the level feature density corresponding to each level, calculate the similarity between the code tree of the component to be detected and the code tree of the standard component.
[0054] Combined with the second possible implementation manner of the second aspect, the embodiment of the present application provides a third possible implementation manner of the second aspect. When the comparison module is used to calculate the level feature density of each level in the code tree of the component to be detected, it is specifically used for:
[0055] For each node in the i-th level of the code tree of the component to be detected, determine the parent-child relationship feature between the node and its parent node according to the node type feature of the node and the node type feature of its parent node, so as to obtain the parent-child relationship feature between each node in the i-th level and its respective parent node; wherein, the value range of i is from 1 to L, and there are L layers in the code tree of the component to be detected except the root node; when i takes 1, each node in the i-th level is the child node of the root node;
[0056] Aggregate the parent-child relationship features between each node in the i-th level and its respective parent node, and perform a hash calculation on the aggregation result to obtain the fingerprint of the i-th level;
[0057] Calculate the Shannon entropy of the node type of this node based on the probabilities of all child nodes of this node for each child node's node type.
[0058] Calculate the Shannon entropy of the node type of each ancestor node of this node respectively, and calculate the depth-aware ancestor entropy of this node based on the Shannon entropy of the node type of each ancestor node of this node.
[0059] Calculate the smooth entropy of the adjacent hierarchical structure relationship corresponding to this node according to the fingerprint of the i-th layer where this node is located and the fingerprint of the (i - 1)-th layer where the parent node of this node is located.
[0060] Calculate the fusion comprehensive entropy of this node based on the Shannon entropy of the node type of this node, the depth-aware ancestor entropy of this node, the smooth entropy of the adjacent hierarchical structure relationship corresponding to this node, and a preset weight coefficient.
[0061] For each node in the i-th layer of the code tree of the component to be detected, calculate the first ratio of the fusion comprehensive entropy of this node to the sum value of the fusion comprehensive entropies of each node in the i-th layer, and calculate the sampling probability of this node according to this first ratio and a preset first attenuation coefficient.
[0062] Select some nodes from each node in the i-th layer as target nodes according to the sampling probabilities of each node in the i-th layer of the code tree of the component to be detected.
[0063] Calculate the initial layer feature density of the i-th layer according to the fingerprint of the i-th layer, the absolute value of the set of node type features of all target nodes in the i-th layer, and the maximum value among the absolute values corresponding to each layer of the code tree of the component to be detected; where the absolute value corresponding to a layer is the absolute value of the set of node type features of all target nodes within this layer.
[0064] Normalize the initial layer feature density of the i-th layer to obtain the layer feature density of the i-th layer.
[0065] Combined with the third possible implementation manner of the second aspect, the embodiments of the present application provide a fourth possible implementation manner of the second aspect. Among them, when the comparison module is used to normalize the initial layer feature density of the i-th layer to obtain the layer feature density of the i-th layer, it specifically is used for:
[0066] Calculate the second ratio of the initial layer feature density of the i-th layer to the sum value of the initial layer feature densities of each layer in the code tree of the component to be detected.
[0067] Calculate the hierarchical feature density of the i-th layer according to the second ratio and the preset second attenuation coefficient.
[0068] Combined with the third possible implementation manner of the second aspect, the embodiment of the present application provides a fifth possible implementation manner of the second aspect. Among them, when the comparison module is used to calculate the similarity degree of the same layer in the code tree of the component to be detected and the code tree of the standard component respectively, it is specifically used for:
[0069] For each node in the i-th layer of the code tree of the component to be detected, perform hash calculation on the node type feature of this node for a preset number of times to obtain the hash value of this node, and perform the preset number of hash calculations on the node type features of each child node of this node to obtain the hash values of each child node, and calculate the sum value of the hash value of this node and the hash values of each child node;
[0070] Perform an exclusive OR calculation on the sum value of the hash value of this node and the hash values of each child node, and the fingerprint of the relationship feature between this node and its parent node, to obtain the matrix corresponding to this node; where multiple element values are included in the matrix;
[0071] Determine the minimum element value in the matrix corresponding to this node;
[0072] After obtaining the minimum element value in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, and use this minimum value as the first MinHash signature corresponding to the i-th layer of the code tree of the component to be detected;
[0073] For the i-th layer, calculate the similarity degree of the i-th layer in the code tree of the component to be detected and the code tree of the standard component according to the first MinHash signature corresponding to the i-th layer of the code tree of the component to be detected and the second MinHash signature corresponding to the i-th layer of the code tree of the standard component.
[0074] Combined with the second aspect, the embodiment of the present application provides a sixth possible implementation manner of the second aspect. Among them, the device further includes:
[0075] A recursion module, used to recursively process the component code file of the component to be detected, traverse the component code file, and obtain the file hierarchical architecture of the component code file;
[0076] A parsing module, used to parse the file hierarchical architecture to obtain each sub-file in the file hierarchical architecture;
[0077] A decompilation module, for each of the sub-files, if the source code contained in the sub-file is bytecode, decompile the bytecode to obtain the source code of the sub-file;
[0078] A second determination module, for determining the source code of all the sub-files in the file hierarchy as the source code of the component to be detected.
[0079] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in any possible implementation manner of the first aspect described above are executed.
[0080] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps in any possible implementation manner of the first aspect described above are executed.
[0081] The component security detection method, device, electronic device, and readable storage medium provided by the embodiments of the present application. In this embodiment, since the tree similarity comparison is performed between the code tree of the component to be detected and the code tree of the standard component, and the code tree of the component to be detected is generated from the source code of the component to be detected, and the code tree of the standard component is generated from the source code of the standard component. Then, if the similarity between the code tree of the component to be detected and the code tree of the standard component is low, it indicates that other code (such as attack code) has been inserted into the source code of the component to be detected. It can be seen that through the method of this embodiment, it is possible to detect whether other code (such as attack code) has been inserted into the source code of the component to be detected, which is beneficial to improving the accuracy of component security detection. At the same time, in this embodiment, by detecting the source code of the component to be detected, there is no need to pay attention to the file list of the component to be detected, and it is not affected by whether each component recorded in the file list has been publicly confirmed to have a security vulnerability. Compared with the method of detecting the file list of the component to be detected, the detection method of this embodiment can not only detect the components that have been publicly confirmed to have security vulnerabilities, but also detect the components that actually have security vulnerabilities but have not been publicly disclosed, thereby improving the accuracy of component security detection.
[0082] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings
[0083] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0084] Figure 1 A flow chart of a component safety detection method provided by an embodiment of the present application is shown;
[0085] Figure 2 A flowchart of another component safety detection method provided by an embodiment of the present application is shown;
[0086] Figure 3 A schematic diagram of the structure of a component safety detection device provided in an embodiment of the present application is shown;
[0087] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0088] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0089] At present, when testing the security of components, the main method is to test the file list that comes with the software application. The file list contains the instructions of each component in the software application. The component instructions record whether the component has been publicly confirmed to have a security vulnerability. In other words, the prior art can only confirm whether the software application contains components that have been publicly confirmed to have security vulnerabilities by testing the instructions of each component. Some components (such as component A) actually have security vulnerabilities, but have not been disclosed, that is, the instructions do not record that component A has been publicly confirmed to have a security vulnerability. At this time, the detection method of the prior art cannot identify that component A has a security vulnerability, resulting in a low accuracy rate in component security detection.
[0090] Furthermore, since the detection method in the prior art is to detect the specifications of each component and cannot detect the code of the component, and there are attack codes inserted into some normal components, the components mixed with attack codes cannot be detected by the detection method in the prior art, which will also lead to low accuracy of component security detection.
[0091] In view of the above problems, based on this, the embodiments of the present application provide a component security detection method, device, electronic device and readable storage medium to improve the accuracy of component security detection, which will be described below through embodiments.
[0092] To facilitate the understanding of this embodiment, first, a component security detection method disclosed in the embodiments of the present application will be introduced in detail. As Figure 1 shown, the component security detection method includes the following steps S101 - S103:
[0093] S101: Perform static analysis on the source code of the component to be detected, and construct a code tree of the component to be detected for the component to be detected.
[0094] In this embodiment, the component to be detected can be an existing component in a software application or a newly added component in a software application. The timing for performing security detection on the component to be detected can be to perform security detection on each component in the software application before deploying the software application, or to perform security detection on the newly added component when a new component is added to the software application.
[0095] In a possible implementation manner, before executing step S101, as Figure 2 shown, the following steps S1001 - S1004 can also be executed:
[0096] S1001: Recursively process the component code file of the component to be detected, traverse the component code file, and obtain the file - level architecture of the component code file.
[0097] S1002: Parse the file - level architecture to obtain each sub - file in the file - level architecture.
[0098] S1003: For each sub - file, if the source code contained in the sub - file is bytecode, decompile the bytecode to obtain the source code of the sub - file.
[0099] S1004: Determine the source code of all sub - files in the file - level architecture as the source code of the component to be detected.
[0100] In step S1001, the component code file contains Python code.
[0101] In step S1003, if the source code included in the sub-file is Bytecode bytecode, decompile the Bytecode bytecode to obtain the source code of the sub-file. If the source code included in the sub-file is source code, do not process the sub-file.
[0102] In a possible implementation manner, when executing step S101, it can be specifically executed according to the following steps S1011 - S1012:
[0103] S1011: Input the source code of the component to be detected into the ast module (Abstract Syntax Tree), and parse the source code of the component to be detected into an abstract syntax tree through the ast module; wherein, the abstract syntax tree contains multiple nodes, and each node corresponds to its own node type.
[0104] S1012: Delete the nodes with the target noise type in the abstract syntax tree from the abstract syntax tree to obtain the code tree of the component to be detected of the component to be detected; wherein, the target noise types include: expression nodes, string nodes, number nodes, name constant nodes, placeholder nodes.
[0105] In this embodiment, after obtaining the abstract syntax tree, considering that the code in the abstract syntax tree is up to thousands of lines, for component information detection, such a complex component code tree is not required to determine the specific information of the component. A large amount of redundant information will occupy a large amount of computing resources when comparing with the existing data, and for large projects, this will greatly increase the load on the database. For the above reasons, this embodiment optimizes the abstract syntax tree to retain the information useful for identifying components to the greatest extent, and obtains the optimized code tree of the component to be detected.
[0106] Among them, in the process of optimizing the abstract syntax tree, the structural feature type nodes are optimally retained, such as Import (import node), ClassDef (class definition node), FunctionDef (function definition node); the behavioral feature type nodes, such as Call (call node), Attribute (attribute node), Assign (assignment node); the metadata features, such as version (version number field), author (author field) nodes as the optimized code tree of the component to be detected.
[0107] Deleted target noise type nodes such as Expr (expression node), Str (string node), Num (number node), NameConstant (name constant node), and Pass (placeholder node). At the same time, we selected three control flow nodes, namely Try (exception handling node), While (conditional loop node), and For (range loop node), as alternatives. When the existing node types cannot determine the component information, the accuracy of component recognition is ensured through the comparison of alternative nodes.
[0108] S102: Compare the tree similarity between the component code tree to be detected and the standard component code tree to obtain the similarity between the component code tree to be detected and the standard component code tree; among them, the standard component code tree is constructed by performing static analysis on the source code of the standard component; the standard component is the standard component of the component to be detected.
[0109] In this embodiment, the component code tree to be detected is the code tree of the component to be detected, and the standard component code tree is the code tree of the standard component, where the standard component is the standard component of the component to be detected. The generation process of generating the standard component code tree based on the standard component is the same as the generation process of generating the component code tree to be detected based on the component to be detected.
[0110] In a possible implementation manner, when executing step S102, it can be specifically executed according to the following steps S1021 - S1023:
[0111] S1021: For each level in the component code tree to be detected, calculate the level feature density of this level.
[0112] S1022: Calculate the similarity degree of the same level in the component code tree to be detected and the standard component code tree respectively, so as to obtain the similarity degree of each level in the component code tree to be detected and the standard component code tree.
[0113] S1023: Calculate the similarity between the component code tree to be detected and the standard component code tree according to the similarity degree corresponding to each level in the component code tree to be detected and the standard component code tree, and the level feature density corresponding to each level.
[0114] When executing step S1021, it can be specifically executed according to the following steps S10210 - S10219:
[0115] S10210: For each node in the i-th level of the code tree of the component to be detected, determine the parent-child relationship feature between the node and its parent node according to the node type feature of the node and the node type feature of the parent node of the node, so as to obtain the parent-child relationship feature between each node in the i-th level and its respective parent node; where, the value range of i is from 1 to L, and there are L layers in the code tree of the component to be detected except the root node; when i takes 1, each node in the i-th level is a child node of the root node.
[0116] In this embodiment, there are L + 1 layers of nodes in the code tree of the component to be detected including the root node, where, l0 represents the root node, l1 represents the first layer of nodes, l i represents the i-th layer of nodes, l L represents the outermost layer of nodes.
[0117] S10211: Aggregate the parent-child relationship features between each node in the i-th level and its respective parent node, and perform a hash calculation on the aggregation result to obtain the fingerprint of the i-th level.
[0118] For each node n in the i-th level (l i ) of the code tree of the component to be detected, calculate the fingerprint of the i-th level through the following formula:
[0119]
[0120] where, type(n) represents the node type feature of node n; type(parent(n)) represents the node type feature of the parent node of node n; type(n)||type(parent(n)) represents the parent-child relationship feature between node n and its parent node in the i-th level.
[0121] represents the aggregation result of aggregating the parent-child relationship features between each node in the i-th level and its respective parent node. Fingerprint(l i ) represents the fingerprint of the i-th level.
[0122] S10212: Calculate the node type Shannon entropy of the node according to the probability of the node type of each child node of the node among all the child nodes of the node.
[0123] In this embodiment, calculate the node type Shannon entropy of node n through the following formula:
[0124] H child (n) = -∑ c∈Children(n) P(c.type)logP(c.type)
[0125] Among them, c represents the child node of node n, and Children(n) represents the set of all child nodes of node n; P(c.type) represents the probability of the node type of the child node c of node n among all child nodes of node n; H child (n) represents the Shannon entropy of the node type of node n.
[0126] S10213: Calculate the Shannon entropy of the node type of each ancestor node of this node respectively, and calculate the depth-aware ancestor entropy of this node according to the Shannon entropy of the node type of each ancestor node of this node.
[0127] In this embodiment, the depth-aware ancestor entropy of node n is calculated by the following formula:
[0128]
[0129] Among them, d represents the number of ancestor nodes of node n; z represents the z-th level ancestor node of node n; Ancestor(n,z) represents the z-th level ancestor node of node n; H child (Ancestor(n,z)) represents the Shannon entropy of the node type of the z-th level ancestor node of node n; H ancestor (n) represents the depth-aware ancestor entropy of node n.
[0130] S10214: Calculate the smooth entropy of the adjacent hierarchical structure relationship corresponding to this node according to the fingerprint of the i-th level where this node is located and the fingerprint of the (i - 1)-th level where the parent node of this node is located.
[0131] In this embodiment, the smooth entropy of the adjacent hierarchical structure relationship corresponding to node n is calculated by the following formula:
[0132] H fingerprint (n) = Jaccard(Fingerprint(l i ), Fingerprint(l i-1 ))
[0133] Among them, Fingerprint(l i ) represents the fingerprint of the i-th level where node n is located; Fingerprint(l i-1 ) represents the fingerprint of the (i - 1)-th level where the parent node of node n is located; H fingerprint (n) represents the smooth entropy of the adjacent hierarchical structure relationship corresponding to the node.
[0134] S10215: Calculate the fusion comprehensive entropy of this node based on the Shannon entropy of the node type of this node, the depth-aware ancestor entropy of this node, the smooth entropy of the adjacent hierarchical structure relationship corresponding to this node, and a preset weight coefficient.
[0135] In this embodiment, the combined entropy H(n) of node n is calculated by the following formula:
[0136] H(n) = αH child (n) + βH ancestor (n) + γH fingerprint (n)
[0137] where α, β, and γ are preset weight coefficients, and α + β + γ = 1. In a preferred example, α = 0.5, β = 0.3, and γ = 0.2.
[0138] S10216: For each node in the i-th level of the code tree of the component to be detected, calculate the first ratio of the combined entropy of this node to the sum of the combined entropies of each node in the i-th level, and calculate the sampling probability of this node according to this first ratio and a preset first attenuation coefficient.
[0139] In this embodiment, the sampling probability P(n) of node n is calculated by the following formula:
[0140]
[0141] where represents the sum of the combined entropies of each node m in the i-th level; σ is the first attenuation coefficient, which controls the attenuation rate.
[0142] S10217: According to the sampling probabilities of each node in the i-th level of the code tree of the component to be detected, select some nodes from each node in the i-th level as target nodes.
[0143] In this embodiment, by setting a threshold for the sampling probability, according to the sampling probability of each node in the i-th level of the code tree of the component to be detected, the nodes with a sampling probability greater than the sampling probability threshold are selected as target nodes. It is also possible to set a threshold for the sampling quantity, and according to the sampling probability of each node in the i-th level of the code tree of the component to be detected, the top sampling quantity threshold nodes with the highest sampling probability are selected as target nodes.
[0144] S10218: According to the fingerprint of the i-th level, as well as the absolute value of the set of node type features of all target nodes in the i-th level and the maximum value among the absolute values corresponding to each level of the code tree of the component to be detected, calculate the initial level feature density of the i-th level; where the absolute value corresponding to a level is the absolute value of the set of node type features of all target nodes within that level.
[0145] In this embodiment, the initial level feature density ρ of the i-th level is calculated by the following formula i :
[0146]
[0147] Among them, S(l i ) represents the set of node type features of all target nodes in the i-th layer; S(l j ) represents the absolute value corresponding to each layer of the code tree of the component to be detected.
[0148] S10219: Normalize the initial layer feature density of the i-th layer to obtain the layer feature density of the i-th layer.
[0149] When performing step S10219, it can be specifically performed according to the following steps:
[0150] Calculate the second ratio of the sum of the initial layer feature density of the i-th layer and the initial layer feature densities of each layer in the code tree of the component to be detected;
[0151] Calculate the layer feature density of the i-th layer according to the second ratio and a preset second attenuation coefficient.
[0152] In this embodiment, the layer feature density of the i-th layer is calculated by the following formula:
[0153]
[0154] Among them, ρ i represents the initial layer feature density of the i-th layer; represents the sum of the initial layer feature densities of each layer in the code tree of the component to be detected; λ is the second attenuation coefficient; ω i is the layer feature density of the i-th layer. In a preferred example, λ = 0.5.
[0155] When performing step S1022, it can be specifically performed according to the following steps S10221 - S10225:
[0156] S10221: For each node in the i-th layer of the code tree of the component to be detected, perform hash calculation on the node type feature of the node for a preset number of times to obtain the hash value of the node, and perform hash calculation on the node type feature of each child node of the node for a preset number of times to obtain the hash values of each child node, and calculate the sum of the hash value of the node and the hash values of each child node.
[0157] In this embodiment, the sum of the hash value of node n and the hash values of each child node is calculated by the following formula:
[0158] ∑ x∈S(n) hash K (x)
[0159] S(n) = {type(n)} ∪ {type(c) | c ∈ children(n)}
[0160] Where type(n) represents the node type feature of node n; children(n) represents the set of children nodes of node n; type(c) represents the node type feature of child c of node n; K represents a preset number of times. In a preferred example, K = 128.
[0161] S10222: Perform an exclusive OR calculation on the sum of the hash values of the node and the hash values of each child node, and the fingerprint of the relationship feature between the node and its parent node, to obtain the matrix corresponding to the node; where the matrix contains multiple element values.
[0162] In this embodiment, the matrix h' corresponding to node n is calculated by the following formula K (x):
[0163] h' K (x) = (∑ x∈S(n) hash K (x)) ⊕ Hash(Fingerprint(l i ))
[0164] S10223: Determine the minimum element value in the matrix corresponding to the node.
[0165] S10224: After obtaining the minimum element value in the matrix corresponding to each node in the i-th layer of the component code tree to be detected, determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the component code tree to be detected, and use this minimum value as the first MinHash signature corresponding to the i-th level of the component code tree to be detected.
[0166] In this embodiment, the first MinHash signature is calculated by the following formula
[0167]
[0168] Where min x∈S(n) (h' K (x)) represents the minimum element value in the matrix corresponding to each node in the i-th layer; is to determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the component code tree to be detected; H1(l i )[k] represents the first MinHash signature.
[0169] S10225: For the i-th level, calculate the similarity between the code tree of the component to be detected and the code tree of the standard component at the i-th level according to the first MinHash signature corresponding to the i-th level of the code tree of the component to be detected and the second MinHash signature corresponding to the i-th level of the code tree of the standard component.
[0170] In this embodiment, the similarity between the code tree of the component to be detected and the code tree of the standard component at the i-th level is calculated by the following formula:
[0171]
[0172] where H2(l i )[k] represents the second MinHash signature corresponding to the i-th level of the code tree of the standard component; Sim i represents the similarity between the code tree of the component to be detected and the code tree of the standard component at the i-th level.
[0173] In this embodiment, the calculation method of the second MinHash signature corresponding to the i-th level of the code tree of the standard component is the same as the calculation method of the first MinHash signature corresponding to the i-th level of the code tree of the component to be detected.
[0174] When executing step S1023, the similarity Sim between the code tree of the component to be detected and the code tree of the standard component can be specifically calculated by the following formula global :
[0175]
[0176] where ω i is the layer feature density corresponding to the i-th level; Sim i represents the similarity between the code tree of the component to be detected and the code tree of the standard component at the i-th level.
[0177] S103: Determine the security of the component to be detected based on the similarity.
[0178] In this embodiment, if the similarity between the code tree of the component to be detected and the code tree of the standard component is higher (for example, greater than the preset similarity threshold), it indicates that the component to be detected is more secure. If the similarity between the code tree of the component to be detected and the code tree of the standard component is lower (for example, less than or equal to the preset similarity threshold), it indicates that the component to be detected is more dangerous.
[0179] Based on the same technical concept, the embodiment of the present application also provides a component security detection device, as shown in Figure 3 The device includes:
[0180] A construction module 301 for performing static analysis on the source code of a component to be detected and constructing a code tree of the component to be detected for the component to be detected;
[0181] A comparison module 302 for comparing the similarity of trees between the code tree of the component to be detected and the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree; wherein, the standard component code tree is constructed by performing static analysis on the source code of the standard component; the standard component is the standard component of the component to be detected;
[0182] A first determination module 303 for determining the security of the component to be detected based on the similarity.
[0183] Optionally, when the construction module 301 is used to perform static analysis on the source code of the component to be detected and construct the code tree of the component to be detected for the component to be detected, it is specifically used for:
[0184] Input the source code of the component to be detected into the ast module, and parse the source code of the component to be detected into an abstract syntax tree through the ast module; wherein, the abstract syntax tree contains multiple nodes, and each node corresponds to its own node type;
[0185] Delete the nodes with the node type of the target noise type from the abstract syntax tree to obtain the code tree of the component to be detected for the component to be detected; wherein, the target noise types include: expression nodes, string nodes, number nodes, name constant nodes, placeholder nodes.
[0186] Optionally, when the comparison module 302 is used to compare the similarity of trees between the code tree of the component to be detected and the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree, it is specifically used for:
[0187] For each level in the code tree of the component to be detected, calculate the level feature density of this level;
[0188] Calculate the similarity degree of the same level in the code tree of the component to be detected and the standard component code tree respectively to obtain the similarity degree of each level in the code tree of the component to be detected and the standard component code tree;
[0189] Calculate the similarity between the code tree of the component to be detected and the standard component code tree according to the similarity degree corresponding to each level in the code tree of the component to be detected and the standard component code tree, and the level feature density corresponding to each level.
[0190] Optionally, when the comparison module 302 is used to calculate the hierarchical feature density of each level in the code tree of the component to be detected, it is specifically used for:
[0191] For each node in the i-th level of the code tree of the component to be detected, determine the parent-child relationship feature between the node and its parent node according to the node type feature of the node and the node type feature of the parent node of the node, so as to obtain the parent-child relationship feature between each node in the i-th level and its respective parent node; where the value range of i is from 1 to L, and there are L layers in the code tree of the component to be detected except the root node; when i takes 1, each node in the i-th level is a child node of the root node;
[0192] Aggregate the parent-child relationship features between each node in the i-th level and its respective parent node, and perform a hash calculation on the aggregation result to obtain the fingerprint of the i-th level;
[0193] Calculate the node type Shannon entropy of the node according to the probability of the node type of each child node of the node among all the child nodes of the node;
[0194] Calculate the node type Shannon entropy of each ancestor node of the node respectively, and calculate the depth-aware ancestor entropy of the node according to the node type Shannon entropy of each ancestor node of the node;
[0195] Calculate the adjacent hierarchical structure relationship smooth entropy corresponding to the node according to the fingerprint of the i-th level where the node is located and the fingerprint of the (i - 1)-th level where the parent node of the node is located;
[0196] Calculate the fusion comprehensive entropy of the node based on the node type Shannon entropy of the node, the depth-aware ancestor entropy of the node, the adjacent hierarchical structure relationship smooth entropy corresponding to the node, and a preset weight coefficient;
[0197] For each node in the i-th level of the code tree of the component to be detected, calculate the first ratio of the sum value of the fusion comprehensive entropy of the node and the fusion comprehensive entropy of each node in the i-th level, and calculate the sampling probability of the node according to the first ratio and a preset first attenuation coefficient;
[0198] Select some nodes from each node in the i-th level of the code tree of the component to be detected as target nodes according to the sampling probability of each node in the i-th level of the code tree of the component to be detected;
[0199] Calculate the initial layer feature density of the i-th layer based on the fingerprint of the i-th layer, the absolute value of the set of node type features of all target nodes in the i-th layer, and the maximum value among the absolute values corresponding to each layer of the code tree of the component to be detected; wherein, the absolute value corresponding to a layer is the absolute value of the set of node type features of all target nodes within that layer.
[0200] Normalize the initial layer feature density of the i-th layer to obtain the layer feature density of the i-th layer.
[0201] Optionally, when the comparison module 302 is used to normalize the initial layer feature density of the i-th layer to obtain the layer feature density of the i-th layer, it is specifically used for:
[0202] Calculate the second ratio of the initial layer feature density of the i-th layer to the sum value of the initial layer feature densities of each layer in the code tree of the component to be detected.
[0203] Calculate the layer feature density of the i-th layer according to the second ratio and a preset second attenuation coefficient.
[0204] Optionally, when the comparison module 302 is used to calculate the similarity degree of the same layer in the code tree of the component to be detected and the standard component code tree respectively, it is specifically used for:
[0205] For each node in the i-th layer of the code tree of the component to be detected, perform hash calculation on the node type feature of this node a preset number of times to obtain the hash value of this node, and perform the preset number of hash calculations on the node type features of each child node of this node to obtain the hash values of each child node, and calculate the sum value of the hash value of this node and the hash values of each child node.
[0206] Perform exclusive OR calculation on the sum value of the hash value of this node and the hash values of each child node, and the fingerprint of the relationship feature between this node and its parent node to obtain the matrix corresponding to this node; wherein, the matrix contains multiple element values.
[0207] Determine the minimum element value in the matrix corresponding to this node.
[0208] After obtaining the minimum element value in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the code tree of the component to be detected, and use this minimum value as the first MinHash signature corresponding to the i-th layer of the code tree of the component to be detected.
[0209] For the i-th level, calculate the similarity degree of the i-th level in the component code tree to be detected and the i-th level in the standard component code tree according to the first MinHash signature corresponding to the i-th level of the component code tree to be detected and the second MinHash signature corresponding to the i-th level of the standard component code tree.
[0210] Optionally, the device further includes:
[0211] A recursive module, configured to recursively process the component code file of the component to be detected, traverse the component code file, and obtain the file-level architecture of the component code file;
[0212] A parsing module, configured to parse the file-level architecture to obtain each sub-file in the file-level architecture;
[0213] A decompiling module, configured to, for each of the sub-files, if the source code included in the sub-file is bytecode, decompile the bytecode to obtain the source code of the sub-file;
[0214] A second determination module, configured to determine the source code of all the sub-files in the file-level architecture as the source code of the component to be detected.
[0215] Figure 4 The figure is a schematic structural diagram of an electronic device provided in an embodiment of the present application, including: a processor 401, a memory 402, and a bus 403. The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device runs the above information processing method, the processor 401 communicates with the memory 402 through the bus 403, and the processor 401 executes the machine-readable instructions to execute the method steps in the first embodiment.
[0216] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the method steps in the first embodiment are executed.
[0217] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described device, electronic device, and computer-readable storage medium can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0218] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices, electronic devices, and computer-readable storage media can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0219] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0220] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0221] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program codes.
[0222] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims described above.
Claims
1. A component safety detection method, characterized in that: The method comprises: Perform static analysis on the source code of the component to be detected, and construct a component code tree of the component to be detected; Performing a tree similarity comparison between the code tree of the component to be detected and the code tree of the standard component to obtain the similarity between the code tree of the component to be detected and the code tree of the standard component; wherein the code tree of the standard component is constructed by statically analyzing the source code of the standard component; and the standard component is the standard component of the component to be detected; The safety of the component to be detected is determined based on the similarity.
2. The method according to claim 1, characterized in that: The static analysis of the source code of the component to be detected is performed to construct a component code tree of the component to be detected, including: Input the source code of the component to be detected into an ast module, and parse the source code of the component to be detected into an abstract syntax tree through the ast module; wherein the abstract syntax tree contains multiple nodes, each node corresponding to a respective node type; Nodes whose node types in the abstract syntax tree are target noise types are deleted from the abstract syntax tree to obtain a code tree of the component to be detected; wherein the target noise types include: expression nodes, string nodes, numeric nodes, name constant nodes, and placeholder nodes.
3. The method according to claim 1, characterized in that: The comparing the tree similarity between the code tree of the component to be detected and the standard component code tree to obtain the similarity between the code tree of the component to be detected and the standard component code tree includes: For each level in the code tree of the component to be detected, calculating the level feature density of the level; Calculating the similarity between the code tree of the component to be detected and the same level in the standard component code tree respectively, so as to obtain the similarity between the code tree of the component to be detected and each level in the standard component code tree; The similarity between the component code tree to be detected and the standard component code tree is calculated according to the similarity corresponding to each level in the component code tree to be detected and the standard component code tree, and the level feature density corresponding to each level.
4. The method according to claim 3, characterized in that: The step of calculating the hierarchical feature density of each level in the code tree of the component to be detected includes: For each node in the i-th level in the code tree of the component to be detected, determine the parent-child relationship feature between the node and the parent node of the node according to the node type feature of the node and the node type feature of the parent node of the node, so as to obtain the parent-child relationship feature between each node in the i-th level and its respective parent node; wherein the value range of i is from 1 to L, and there are L layers in the code tree of the component to be detected except the root node; when i is 1, each node in the i-th level is a child node of the root node; Aggregate the parent-child relationship features between each node in the i-th level and its respective parent node, and perform hash calculation on the aggregation results to obtain the fingerprint of the i-th level; According to the probability of the node type of each child node of the node in all the child nodes of the node, the node type Shannon entropy of the node is calculated; Calculate the node type Shannon entropy of each ancestor node of the node respectively, and calculate the depth-aware ancestor entropy of the node according to the node type Shannon entropy of each ancestor node of the node; According to the fingerprint of the i-th level where the node is located and the fingerprint of the i-1-th level where the parent node of the node is located, the smooth entropy of the adjacent hierarchical structure relationship corresponding to the node is calculated; Calculate the fusion comprehensive entropy of the node based on the node type Shannon entropy of the node, the deep perception ancestor entropy of the node, the smooth entropy of the adjacent hierarchical structure relationship corresponding to the node, and the preset weight coefficient; For each node in the i-th level of the code tree of the component to be detected, calculate a first ratio of the fused comprehensive entropy of the node to the sum of the fused comprehensive entropies of each node in the i-th level, and calculate the sampling probability of the node according to the first ratio and a preset first attenuation coefficient; According to the sampling probability of each node in the i-th level of the code tree of the component to be detected, select some nodes from each node in the i-th level as target nodes; Calculate the initial hierarchical feature density of the i-th level according to the fingerprint of the i-th level, the absolute value of the set of node type features of all target nodes in the i-th level, and the maximum value of the absolute values corresponding to each level of the code tree of the component to be detected; wherein the absolute value corresponding to the level is the absolute value of the set of node type features of all target nodes in the level; The initial level feature density of the i-th level is normalized to obtain the level feature density of the i-th level.
5. The method according to claim 4, characterized in that: The step of normalizing the initial level feature density of the i-th level to obtain the level feature density of the i-th level includes: Calculating a second ratio of an initial level feature density of the i-th level to a sum of the initial level feature densities of each level in the code tree of the component to be detected; The level feature density of the i-th level is calculated according to the second ratio and a preset second attenuation coefficient.
6. The method according to claim 4, characterized in that: The respectively calculating the similarity between the code tree of the component to be detected and the code tree of the standard component at the same level includes: For each node in the i-th level of the code tree of the component to be detected, perform a hash calculation on the node type feature of the node for a preset number of times to obtain a hash value of the node, and perform a hash calculation on the node type feature of each child node of the node for the preset number of times to obtain a hash value of each child node, and calculate the sum of the hash value of the node and the hash values of each child node; Performing XOR calculation on the sum of the hash value of the node and the hash values of each child node, and the fingerprint of the relationship feature between the node and the parent node of the node, to obtain a matrix corresponding to the node; wherein the matrix contains multiple element values; Determine the minimum element value in the matrix corresponding to the node; After obtaining the minimum element value in the matrix corresponding to each node in the i-th layer of the component code tree to be detected, determine the minimum value from the minimum element values in the matrix corresponding to each node in the i-th layer of the component code tree to be detected, and use the minimum value as the first MinHash signature corresponding to the i-th level of the component code tree to be detected; For the i-th level, based on the first MinHash signature corresponding to the i-th level of the component code tree to be detected and the second MinHash signature corresponding to the i-th level of the standard component code tree, the similarity between the component code tree to be detected and the i-th level in the standard component code tree is calculated.
7. The method according to claim 1, characterized in that: Before statically analyzing the source code of the component to be detected and constructing a component code tree of the component to be detected, the method further includes: Recursively traverse the component code files of the component to be detected, and obtain the file hierarchy structure of the component code files; Parsing the file hierarchy structure to obtain each sub-file in the file hierarchy structure; For each of the sub-files, if the source code contained in the sub-file is a bytecode, decompile the bytecode to obtain the source code of the sub-file; The source codes of all the sub-files in the file hierarchy are determined as the source codes of the component to be detected.
8. A component safety detection device, characterized in that: The device comprises: A construction module is used to perform static analysis on the source code of the component to be detected and construct a component code tree of the component to be detected; A comparison module, used for performing a tree similarity comparison between the code tree of the component to be detected and the code tree of the standard component, to obtain the similarity between the code tree of the component to be detected and the code tree of the standard component; wherein the code tree of the standard component is constructed by statically analyzing the source code of the standard component; and the standard component is the standard component of the component to be detected; The first determination module is used to determine the safety of the component to be detected based on the similarity.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the method as described in any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are executed.