Software supply chain security monitoring method and system
By conducting in-depth analysis of the code snippets and binary code of open source software, combined with Spark's big data processing capabilities, identifying and marking open source software components, the problem of inefficient malicious code monitoring in the existing technology is solved, and the security and compliance management of the software are realized.
Patent Information
- Application Number
- CN202510344922.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology is difficult to efficiently identify and monitor malicious code in open source software, resulting in an increase in security incidents and inefficient analysis, which cannot meet the current information security requirements.
By obtaining code snippets of open source software, establishing a Spark analysis model, performing word vector representation and feature extraction, combining multi-dimensional feature analysis of binary code, identifying open source software components, and label printing based on machine learning.
It improves the accuracy of open source software component detection, ensures software compliance and security, and promotes intelligent management and tracking of open source code.
Smart Images

Figure CN120234805A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a software supply chain security monitoring method and system. Background Art
[0002] New technologies such as big data of open-source software, cloud computing, and machine learning are leading the development of technology, and most of these new technologies are open-source software. These open-source software have been widely applied to various industries, such as the financial industry, e-commerce, information security, and the education industry. With the increasingly widespread use of open-source software, security incidents caused by the security risks of open-source software are also increasing year by year. Early methods for detecting malicious code mainly included: signature-based detection, integrity-based detection, heuristic and semantic-based detection, etc. However, these detection methods have problems such as over-reliance on signature codes, over-reliance on the experience of analysts, and inability to detect unknown malicious codes. Facing the current large number of malicious codes, the analysis efficiency is low, and it is difficult to meet the current information security requirements. Summary of the Invention
[0003] The purpose of the present invention is to provide a software supply chain security monitoring method and system, which solves the problem of meeting the current information security requirements for monitoring malicious codes.
[0004] To achieve the above purpose, the present invention provides a software supply chain security monitoring method, and the monitoring method includes:
[0005] Obtain code snippets of open-source software;
[0006] Build a code snippet analysis model based on Spark;
[0007] Input the code snippets of the open-source software into the code snippet analysis model to implement component analysis of the open-source software;
[0008] Obtain the binary code of the open-source software;
[0009] Monitor the binary code to implement component analysis of the open-source software;
[0010] Label the code snippets and binary code according to the results of the component analysis of the open-source software.
[0011] Optionally, inputting the code snippets of the open-source software into the code snippet analysis model to implement component analysis of the open-source software includes:
[0012] Preprocess the code snippets of the open-source software using a word vector representation method to form preprocessed code, and the word vector representation method includes GloVe and Word2vec;
[0013] Extract features from the preprocessed code to establish a code snippet feature library, including:
[0014] Combine the tokens in the preprocessed code in a fixed number to form code snippets;
[0015] Calculate the Hash value of the code snippet to obtain the feature of the code snippet;
[0016] Shift a fixed number of tokens to form a new code snippet and calculate the Hash value of the new snippet;
[0017] Determine whether to traverse the preprocessed code;
[0018] In the case of determining that the preprocessed code has been traversed, end the operation;
[0019] In the case of determining that the preprocessed code has not been traversed, perform the operation of calculating the Hash value of the code snippet to obtain the feature of the code snippet;
[0020] Perform open-source software component analysis according to the code snippet feature library.
[0021] Optionally, performing open-source software component analysis according to the code snippet feature library includes:
[0022] Construct an index of the source code to be monitored and use the index to retrieve in HBase;
[0023] Adopt the two-dimensional adjacency list structure generated by the retrieval, use the index sequence of the code to be monitored as the row header, store the index data with the same hash value as the row header element in each row, sort the elements in each row according to the file path, and sort the row headers according to their index numbers in the file to be monitored;
[0024] Analyze and identify software components according to the index and the open-source software distributed index library.
[0025] Optionally, converting the code of the open-source software into binary code to implement open-source software component analysis, including:
[0026] Convert the code of the open-source software into binary code;
[0027] Extract multi-dimensional features of the binary code, where the multi-dimensional features include file attribute features, binary structure features of the file structure layer, byte sequence features of the byte layer, instruction sequence features of the instruction layer, and key API call sequence features of the semantic layer;
[0028] By performing multi-dimensional feature extraction on binary open-source software and binary malicious code, a binary open-source software feature library and a binary malicious code feature library are established, and the open-source components in the binary code are identified based on feature matching of basic block code sequences, feature matching of basic block sub-function names, and feature matching of constants;
[0029] Analyze the security risks summarized in the binary code.
[0030] Optionally, the steps of extracting multi-dimensional features of the binary code include:
[0031] Use Hexview to convert the binary code into hexadecimal code;
[0032] Use an n-gram window to slide to obtain the byte-level features;
[0033] Construct a key system call library, and the key system call library includes key system call API functions;
[0034] Generate a key system call relationship flow graph based on the disassembly result and the key system call library;
[0035] Adopt a depth-first traversal algorithm to traverse the call flow graph to obtain the call sequence of key system calls;
[0036] Use an n-gram window to slide the call sequence to obtain the semantic-level features.
[0037] Optionally, by performing multi-dimensional feature extraction on binary open-source software and binary malicious code, a binary open-source software feature library and a binary malicious code feature library are established, and the open-source components in the binary code are identified based on feature matching of basic block code sequences and feature matching of basic block sub-function names. The feature matching based on basic block code sequences includes:
[0038] By referring to the matrix mapping method and scoring matrix method and ideas for calculating DNA sequence similarity in biology, use two-dimensional array programming to simulate a dot matrix to realize the similarity measurement of executable binary assembly code sequences.
[0039] Extract the assembly instruction operators from two binary assembly code sequences to be compared, and represent them respectively as:
[0040] sequenceA = {A1, A2, A3, …, A i , …, A lengthA}
[0041] sequenceB = {B1, B2, B3, …, B i , …, B lengthA}
[0042] The sliding window value is set to H, and the scoring rule is as follows:
[0043]
[0044] The scoring matrix ScoreMatrix is obtained by comparing sequenceA and sequenceB, and the sequence length of the longest combination of non-overlapping parallel diagonal lines in the scoring matrix ScoreMatrix is obtained by scanning it:
[0045]
[0046] The length of the same sequence appearing in sequenceA and sequenceB is SC. This matrix algorithm can detect continuous identical instruction sequences, effectively eliminate the influence of instruction reordering, and accurately obtain the similarity between the two instruction sequences.
[0047] Optionally, by performing multi-dimensional feature extraction on binary open-source software and binary malicious code, a binary open-source software feature library and a binary malicious code feature library are established, and the open-source components in the binary code are identified based on the feature matching of the basic block code sequence and the feature of the basic block sub-function name. The identification of the open-source components in the binary code by the feature of the basic block sub-function name includes:
[0048] Extract the valid sub-function name from the operand of the call instruction as an important matching feature, and calculate the minimum number of steps required to make two strings the same. The operation method for each step is to modify, add, or delete a character. This minimum number of steps is the edit distance ED.
[0049] Assume basic block A and basic block B. Scan all call instructions in basic blocks A and B respectively, extract the operands of the call instructions, i.e., the sub-function names, filter out the function names that are 8 hexadecimal numbers, and obtain two string sets StrA and StrB, and m > n.
[0050] StrA = (m, StrA1, StrA2…StrA m )
[0051] StrB = (n, StrB1, StrB2…StrB n )
[0052] The similarity between the two string sets is:
[0053]
[0054] Among them, F Similarity(StrA, StrB) is the similarity, StrA is the first string, StrB is the second string, ED is the edit distance, m is the length of the first string, n is the length of the second string, and 1 ≤ j < n, 1 ≤ i < m.
[0055] Optionally, identifying the open-source components in the binary code based on the feature matching of binary code constant features includes:
[0056] Extracting the constant features in the binary code to implement a feature matching method independent of the hardware platform and compiler;
[0057] By extracting the constant features in the binary open-source software, an open-source software feature library is established;
[0058] Using the feature matching method based on the inverted index to implement the analysis of open-source components in the binary code.
[0059] Optionally, analyzing the security risks in the binary code includes:
[0060] By matching the multi-dimensional features of the binary code, the open-source software components and versions in the monitored binary code are analyzed, and through the global open-source software information knowledge base, the security vulnerability risk situation of the open-source software of this version is obtained, thereby realizing security risk analysis;
[0061] Based on the malicious code features in the global open-source software risk feature library, machine learning is used for learning and training to obtain a binary file security risk analysis model and a rule library, and through the model and the rule library, the security risk analysis of the binary file is realized.
[0062] On the other hand, the present invention provides a software supply chain security monitoring system, and the monitoring system includes a processor for executing any one of the above monitoring methods.
[0063] Through the above technical solutions, the present invention provides a software supply chain security monitoring method and system. By obtaining the code fragments of the open-source software and the binary code and combining the Spark big data processing ability, in-depth analysis of the open-source software components is realized. First, through the analysis of the code fragments, the specific components of the open-source software can be identified; secondly, through the monitoring of the binary code, the component information can be further confirmed and supplemented. This process not only improves the detection accuracy of the used components of the open-source software, but also can timely label and indicate the open-source libraries or components it depends on, ensuring the compliance and security of the software, and at the same time promoting the intelligent management and tracking of the open-source code.
[0064] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the accompanying drawings:
[0066] Figure 1 is a flowchart of a monitoring method according to an embodiment of the present invention;
[0067] Figure 2 is a flowchart of component analysis of open-source software according to code snippets according to an embodiment of the present invention;
[0068] Figure 3 is a flowchart of component analysis of open-source software according to a code snippet feature library according to an embodiment of the present invention;
[0069] Figure 4 is a flowchart of component analysis of open-source software according to binary code according to an embodiment of the present invention;
[0070] Figure 5 is a flowchart of feature extraction according to an embodiment of the present invention;
[0071] Figure 6 is a flowchart of byte-layer feature extraction according to an embodiment of the present invention;
[0072] Figure 7 is a flowchart of semantic-layer feature extraction according to an embodiment of the present invention;
[0073] Figure 8 is a flowchart of identifying open-source components in binary code according to an embodiment of the present invention. Specific Embodiments
[0074] The following details the specific embodiments of the embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0075] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant regulations of national laws and regulations. In the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0076] Figure 1FIG. 0 is a flowchart of a monitoring method according to an embodiment of the present invention. In this figure, the monitoring method includes:
[0077] In step S1, code snippets of open-source software are obtained.
[0078] In step S2, a code snippet analysis model is established based on Spark.
[0079] In step S3, the code snippets of the open-source software are input into the code snippet analysis model to implement component analysis of the open-source software.
[0080] In step S4, the binary code of the open-source software is obtained.
[0081] In step S5, the binary code is monitored to implement component analysis of the open-source software.
[0082] In step S6, tags are assigned to the code snippets and the binary code according to the results of the component analysis of the open-source software.
[0083] Compared with the prior art, the present invention can accurately identify various components in open-source software and label them through comprehensive analysis of code snippets and binary code. This helps developers better understand the composition of the software, improve development efficiency, and also facilitates security assessment and compliance checking.
[0084] To facilitate the source code analysis model to analyze the code snippet, the code snippet needs to be preprocessed. Specifically, as Figure 2 shown, it includes the following steps:
[0085] In step S31, a lexical analyzer is used to preprocess the code snippets of the open-source software to form preprocessed code. The lexical analyzer is used to denoise the irrelevant parts of the code, filter comments, header files, and redundant meaningless characters. After obtaining standardized lexical units, fixed-length index units are further obtained to implement the preprocessing of the code snippet.
[0086] In step S32, feature extraction is performed on the preprocessed code to establish a code snippet feature library. By performing feature extraction on the original code and the preprocessed code, a code snippet feature library can be established. Taking the preprocessed code as an example, the method for extracting code snippet features is as follows: First, the tokens in the preprocessed code are combined according to a fixed number to form a code snippet; then, the hash value of the code snippet is calculated, and the feature of the code snippet is obtained; then, the tokens are offset by a fixed number and combined according to a fixed number to form a new code snippet, and the hash value of the new snippet is calculated.
[0087] In step S33, open-source software component analysis is performed according to the code snippet feature library.
[0088] In this embodiment, there are various ways to perform open-source software component analysis according to the code snippet feature library, which are known to those skilled in the art. In one example of the present invention, such as Figure 3 , the following steps may be included:
[0089] In step S331, an open-source software distributed index library is established based on the preprocessed code.
[0090] In step S332, an index of the source code to be monitored is constructed and retrieved using the index in HBase. By constructing an index of the source code to be monitored and storing the index in the HBase database, this step aims to achieve efficient retrieval of large-scale source code. The distributed nature of HBase enables data to be accessed and processed quickly, thus supporting rapid location and query of a large number of code files and improving the retrieval efficiency.
[0091] In step S333, the retrieved two-dimensional adjacency list structure is adopted. Using the index sequence of the code to be monitored as the row header, each row stores index data with the same hash value as the row header element. The elements of each row are sorted according to the file path, and the row headers are sorted according to their index numbers in the file to be monitored. This structure optimizes the data organization method, making subsequent processing more efficient and facilitating the rapid identification and analysis of the association relationships between codes.
[0092] In step S334, software components are analyzed and identified based on the index and the open-source software distributed index library.
[0093] In view of the hardware platform dependence, compiler dependence, etc. existing in binary code analysis, the present invention proposes an automated analysis and identification method by studying and analyzing the multi-dimensional features of binary code. Specifically, such as Figure 4 shown, the following steps are included:
[0094] In step S51, the code of the open-source software is converted into binary code.
[0095] In step S52, multi-dimensional features of the binary code are extracted, and a binary code feature library is established.
[0096] In step S53, the open-source components in the binary code are identified.
[0097] In step S54, the security risks summarized from the binary code are analyzed. The present invention realizes the open-source component analysis of binary code under various hardware platforms, providing key technical support for software supply chain security monitoring.
[0098] In this embodiment, the features of the extracted binary code can be various known to those skilled in the art. In one example of the present invention, as Figure 5 shown, it includes:
[0099] In step S521, extract the file attribute features of the binary code. The file attribute features include the features represented by basic attributes such as file name, capacity, hash value, etc. This feature can be used for the preliminary identification of open-source binary software or binary components. For binary components that can directly identify the version, security risk analysis can be preliminarily carried out based on their vulnerability conditions. In combination with subsequent features, multi-dimensional feature analysis of the binary code can be further carried out to identify security risks such as malicious code.
[0100] In step S522, extract the structural layer features of the binary code. The features of the file structure layer focus on the static structure information of the file. In order to implement functions such as relocation, file search, infection and destruction, and prevent being detected by anti-virus software, malicious code usually modifies the structure of the binary file to achieve its own purpose, such as modifying the entry point to point to a non-standard section, modifying the section name, modifying the import table, etc. These structural features are different from normal files. Therefore, the structural features of the binary file can be used as one of the binary file features to analyze and identify possible malicious code.
[0101] In step S523, extract the byte layer features of the binary code.
[0102] In step S524, extract the instruction layer features of the binary code.
[0103] In step S525, extract the semantic layer features of the binary code.
[0104] In this embodiment, the methods for extracting byte layer features can be various known to those skilled in the art. In one example of the present invention, as Figure 6 shown, it may include the following steps:
[0105] In step S5231, use Hexview to convert the binary code into hexadecimal code.
[0106] In step S5232, use the n-gram window to slide to obtain the byte layer features.
[0107] In this embodiment, the methods for extracting semantic layer features can be various known to those skilled in the art. In one example of the present invention, as Figure 7 shown, it may include the following steps:
[0108] In step S5251, construct a key system call library, and the key system call library includes key system call API functions.
[0109] In step S5252, a critical system call relationship flow graph is generated based on the disassembly result and the critical system call library.
[0110] In step S5253, the call flow graph is traversed using the depth-first traversal algorithm to obtain the call sequence of the critical system calls.
[0111] In step S5254, the n-gram window is used to slide the call sequence to obtain the required features.
[0112] In this embodiment, for the method of identifying open-source components in binary code, there can be various methods known to those skilled in the art. In one example of the present invention, as Figure 7 shown, the following steps may be included:
[0113] In step S531, the second assembly instruction operator corresponding to the first assembly instruction operator in the sequences of the binary code and the target code is extracted according to formulas (1) and (2).
[0114] sequenceA = {A1, A2, A3, …, A i , …, A lengthA}, (1)
[0115] sequenceB = {B1, B2, B3, …, B i , …, B lengthA}, (2)
[0116] where sequenceA represents the binary assembly code sequence A to be compared, sequenceB represents the binary assembly code sequence B to be compared; A i and B i are respectively the i-th assembly instruction operator in sequenceA and sequenceB, and lengthA is the length of the sequences of the binary code and the target code.
[0117] In step S532, a sliding window is set for scoring and comparison. The sliding window is used to slide and select the assembly instruction operators to be compared in the binary assembly code sequence. The scoring rule is:
[0118]
[0119] where P(A i , B i ) is the scoring function.
[0120] In step S533, sequenceA and sequenceB are compared to obtain the scoring matrix ScoreMatrix.
[0121] In step S534, scan the ScoreMatrix to obtain the sequence length of the longest combination of non-overlapping parallel diagonal lines of the scoring matrix:
[0122]
[0123] In step S535, determine the same sequence length according to the sequence length of the longest combination of non-overlapping parallel diagonal lines of the scoring matrix. The present invention can detect continuous identical instruction sequences, effectively eliminate the influence of instruction rearrangement, and accurately obtain the similarity between two instruction sequences. By drawing on the matrix mapping method and scoring matrix method and ideas for calculating the similarity of DNA sequences in biology, a two-dimensional array is used to program and simulate a dot matrix to achieve the similarity measurement of executable binary assembly code sequences.
[0124] In this embodiment, in order to further improve the monitoring efficiency, tags are added to the code fragments and binary codes according to the results of the component analysis of open-source software. Specifically, it includes adding a tag of no security risk to the code with a safe analysis result, and vice versa, adding a tag of security risk. Multiple extracted code feature data are used for security analysis, and a tag is added to the security of the code. Through this tag, the security of the code can be analyzed.
[0125] On the other hand, the present invention provides a software supply chain security monitoring system. The monitoring system includes a processor for executing any one of the above monitoring methods.
[0126] Through the above technical solutions, the present invention provides a software supply chain security monitoring method and system. By obtaining the code fragments and binary codes of open-source software and combining the Spark big data processing ability, in-depth analysis of the components of open-source software is realized. First, through the analysis of code fragments, the specific components of open-source software can be identified; second, through the monitoring of binary codes, the component information can be further confirmed and supplemented. This process not only improves the detection accuracy of the components used in open-source software, but also can add tags in time to indicate the open-source libraries or components it depends on, ensuring the compliance and security of the software, and at the same time promoting the intelligent management and tracking of open-source code.
[0127] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0128] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A software supply chain security monitoring method, characterized in that: The monitoring method comprises: Get code snippets from open source software; Establish a code snippet analysis model based on Spark; Inputting the code snippet of the open source software into the code snippet analysis model to implement component analysis of the open source software; Get binary code of open source software; Monitoring the binary code to implement component analysis of the open source software; The code snippets and binary codes are labeled according to the results of the component analysis of the open source software.
2. The monitoring method according to claim 1, characterized in that: Inputting the code snippet of the open source software into the code snippet analysis model to implement component analysis of the open source software includes: Preprocessing the code snippet of the open source software using a word vector representation method to form a preprocessed code, wherein the word vector representation method includes GloVe and Word2vec; Extracting features from the preprocessed code to establish a code snippet feature library includes: Combining tokens in the preprocessed code according to a fixed number to form a code snippet; Performing hash value calculation on the code snippet to obtain features of the code snippet; Offset a fixed number of tokens to form a new code snippet and calculate the hash value of the new snippet; Determine whether to traverse the preprocessing code; If it is determined that the preprocessing code has been traversed, the operation ends; In the case where it is determined that the preprocessing code has not been traversed, performing a hash value calculation on the code fragment to obtain a feature operation of the code fragment; Open source software component analysis is performed based on the code snippet feature library.
3. The monitoring method according to claim 2, characterized in that: Performing open source software component analysis based on the code snippet feature library includes: Construct an index of source code to be monitored, and use the index to search HBase; The two-dimensional adjacency list structure generated by the retrieval is adopted, and the index sequence of the code to be monitored is used as the row header, and each row stores the index data having the same hash value as the row header element, and the elements of each row are sorted according to the file path, and the row header is sorted according to the index number in the file to be monitored; Software components are analyzed and identified based on the index and the open source software distributed index library.
4. The monitoring method according to claim 1, characterized in that: Converting the code of the open source software into binary code to implement component analysis of the open source software includes: Convert the code of the open source software into binary code; Extracting multi-dimensional features of the binary code, the multi-dimensional features including file attribute features, binary structure features of the file structure layer, byte sequence features of the byte layer, instruction sequence features of the instruction layer, and key API call sequence features of the semantic layer; By performing multi-dimensional feature extraction on binary open source software and binary malicious code, a binary open source software feature library and a binary malicious code feature library are established, and the open source components in the binary code are identified based on feature matching of basic block code sequences, feature matching of basic block sub-function names, and feature matching of constant features; The security risks of the binary code summary are analyzed.
5. The monitoring method according to claim 4, characterized in that: The step of extracting the multi-dimensional features of the binary code comprises: Convert the binary code into hexadecimal code using Hexview; The byte-level features are obtained by sliding an n-gram window; Building a key system call library, wherein the key system call library includes a key system call API function; Generate a key system call relationship flow graph based on the disassembly result and the key system call library; Using a depth-first traversal algorithm to traverse the call flow graph to obtain a call sequence of key system calls; The n-gram window is used to slide the call sequence to obtain semantic layer features.
6. The monitoring method according to claim 4, characterized in that: By performing multi-dimensional feature extraction on binary open source software and binary malicious code, a binary open source software feature library and a binary malicious code feature library are established, and the open source components in the binary code are identified based on feature matching of basic block code sequences and basic block sub-function name features. Feature matching based on basic block code sequences includes: Extract the assembly instruction operators for the two binary assembly code sequences that need to be compared, which are expressed as: sequenceA={A1,A2,A3,…,A i ,…,A lengthA } sequenceB={B1,B2,B3,…,B i ,…,B lengthA } sequenceA represents the binary assembly code sequence A that needs to be compared, and sequenceB represents the binary assembly code sequence B that needs to be compared; A i and B i are the i-th assembly instruction operators in sequenceA and sequenceB respectively, and lengthA is the length of the sequence of binary code and target code; Set a sliding window for scoring and comparison. The sliding window is used to slide and select the assembly instruction operators to be compared in the binary assembly code sequence. The scoring rules are: Compare sequenceA and sequenceB to get the scoring matrix ScoreMatrix, scan ScoreMatrix to get the sequence length of the longest combination of non-overlapping parallel diagonal lines in the scoring matrix: The identical sequence length is determined according to the sequence length of the longest combination of non-overlapping parallel oblique lines of the scoring matrix.
7. The monitoring method according to claim 4, characterized in that: By performing multi-dimensional feature extraction on binary open source software and binary malicious code, a binary open source software feature library and a binary malicious code feature library are established, and the open source components in the binary code are identified based on feature matching of basic block code sequences and basic block sub-function name features. The basic block sub-function name features identify the open source components in the binary code, including: Assume basic blocks A and B, scan all call instructions of basic blocks A and B respectively, extract the operands of call instructions, i.e., sub-function names, filter out function names with 8 hexadecimal numbers, and obtain two string sets StrA and StrB, and m>n; StrA=(m,StrA1,StrA2…StrA m ) StrB=(n,StrB1,StrB2…StrB n ) The similarity between two sets of strings is: Among them, F Similarity (StrA, StrB) is the similarity, StrA is the first string, StrB is the second string, ED is the edit distance, m is the length of the first string, n is the length of the second string, and 1≤j <n,1≤i<m。 8. The monitoring method according to claim 4, characterized in that: Identifying open source components in the binary code based on feature matching of binary code constant features includes: Extracting constant features from the binary code to implement a feature matching method that is independent of the hardware platform and compiler; By extracting constant features from binary open source software, an open source software feature library is established; Use the inverted index-based feature matching method to implement open source component analysis in binary codes.
9. The security analysis method according to claim 4, characterized in that: Analyze security risks in binary code, including: By matching the multi-dimensional features of binary codes, we can analyze the open source software components and versions in the monitored binary codes. Through the global open source software information knowledge base, we can obtain the security vulnerability risk situation of the open source software version, thus achieving security risk analysis. Based on the malicious code features in the global open source software risk feature library, machine learning is used for learning and training to obtain a binary file security risk analysis model and rule library. Through the model and rule library, security risk analysis of binary files is achieved.
10. A monitoring system, characterized in that: The monitoring system comprises: A monitoring module, used to monitor the status of the monitored binary files in the target system; The data transmission module is used to compress the data to be transmitted, divide it into packets and send it to the designated target; Feature extraction module, used to extract some simple features of binary files; The decision module manages the working status of the monitoring module, the data transmission module and the feature extraction module according to the status and output of the monitoring module, the data transmission module and the feature extraction module, and makes analysis decisions on the security risks of the binary file.