Adaptive encryption algorithm dynamic detection method based on abstract syntax tree
By adopting an adaptive encryption algorithm dynamic detection method based on abstract syntax trees, the problem that encryption detection technology cannot adapt to encryption library iteration and lacks context awareness is solved, and efficient and accurate encryption algorithm detection is achieved, which meets the needs of modern development processes.
Patent Information
- Application Number
- CN202510962189.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing encryption detection technologies cannot automatically adapt to encryption library iterations and lack context awareness, resulting in high false negative rates and low detection efficiency, failing to meet the timeliness requirements of modern development processes.
An adaptive encryption algorithm dynamic detection method based on abstract syntax trees is adopted. By constructing an initial encryption rule base, dynamic updates and feature extraction are performed. Combined with full-stack context modeling and distributed processing architecture, efficient detection of encryption algorithms is achieved.
It achieves automatic adaptation to emerging encryption libraries, reduces false negative rates, improves detection efficiency, and can handle tens of thousands of codebases, meeting the timeliness requirements of modern DevSecOps pipelines.
Smart Images

Figure CN120848943A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software security detection technology, and in particular to an adaptive encryption algorithm dynamic detection method based on abstract syntax trees. Background Technology
[0002] Currently, in the field of software security testing, the detection of encryption algorithm usage is a core aspect of ensuring system security. With the rapid development of fintech, cloud services, and the Internet of Things, numerous sensitive data processing systems employ various encryption algorithms to protect data security. However, traditional encryption algorithms (such as RSA and DES) have been gradually phased out by modern standards due to known vulnerabilities, and their continued use poses a significant security risk. Current mainstream detection technologies primarily rely on static code analysis, but these have significant limitations when dealing with large-scale codebases and dynamic language characteristics. While regular expression-based text matching methods are simple to implement, they cannot identify dynamic call paths and alias import scenarios for encryption algorithms, and are completely ineffective against advanced features such as Python decorators and metaclass programming, resulting in a false negative rate exceeding 88%, severely restricting the effectiveness of detection.
[0003] While abstract syntax tree (AST) analysis techniques with fixed rule bases can partially solve the structure recognition problem, they expose fundamental flaws when facing the rapidly iterating cryptographic ecosystem. These methods rely on predefined cryptographic feature rules and cannot automatically adapt to emerging cryptographic libraries (such as Cryptography replacing PyCrypto) and custom wrapper modules; the rule base needs to be manually updated every time the algorithm is upgraded. More critically, existing technologies generally lack context awareness, failing to trace cross-file call chains and class inheritance relationships, resulting in weak business scenario localization capabilities for cryptographic algorithms. For example, in multi-module projects, cryptographic initialization may be completed at the service layer while the actual calls occur at the data layer; traditional solutions struggle to establish a complete call graph. While dynamic runtime analysis can capture the actual execution path, it consumes enormous resources and has insufficient path coverage, requiring hundreds of hours to process tens of thousands of codebases, failing to meet the timeliness requirements of modern DevSecOps pipelines.
[0004] In summary, current encryption detection technologies face three major technological gaps: First, the contradiction between static rules and the rapid evolution of cryptography means existing solutions cannot automatically adapt to encryption library iterations; second, the disconnect between contextual fragmentation and business risk identification, lacking scenario-based analysis of encryption algorithms within business logic; and third, the conflict between scale bottlenecks and modern development processes, with the detection time for tens of thousands of codebases far exceeding the requirements of CI / CD pipelines. These shortcomings expose various highly sensitive fields to serious compliance risks, urgently requiring innovative solutions to overcome these technological bottlenecks. Summary of the Invention
[0005] In view of the above-mentioned problems in the existing technology, the technical problem to be solved by the present invention is: how to establish an encryption rule base with adaptive evolution capability, and to realize high-efficiency encryption algorithm usage detection based on the rule base.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0007] An adaptive encryption algorithm dynamic detection method based on abstract syntax trees includes the following steps:
[0008] S100: Construct an initial encryption rule base. Extract several encryption algorithm feature rules from publicly available cryptography literature using a regularization matching algorithm to form the initial encryption rule base. The encryption algorithm features include encryption function labels and the class to which the encryption function belongs.
[0009] S200: Select a public code repository, which includes source code files or a compressed code repository; preprocess the code repository to obtain a target code repository, where each source code file in the target code repository has an encrypted function tag;
[0010] S300: Dynamically update the initial encryption rule base to obtain an updated encryption rule base;
[0011] S400: Based on the encryption rule update library, the target code library is scanned again to locate the positions of all code source code in the target code library that uses encryption algorithms, and a set of positions is obtained. An encryption function relationship graph is constructed, which consists of all encryption functions called by each code source code. The encryption function relationship graphs of all code source code form a relationship graph set.
[0012] S500: Outputs the location set and relationship graph set to obtain the detection results for all source code that uses encryption algorithms.
[0013] Preferably, the content of the target code library obtained by preprocessing the code library in S200 is as follows:
[0014] First, the source code file is checked; if it is the target source code file, it is kept; otherwise, it is filtered and discarded.
[0015] Secondly, the compressed code repository is checked. First, the directory structure of the code repository is decompressed to a temporary directory. Then, the source files in the temporary directory are checked to see if they are the target source files. If they are, they are kept; otherwise, they are filtered and discarded.
[0016] Finally, the retained target source code files are used to form the target code library.
[0017] Preferably, the content of the updated encryption rule library obtained by dynamically updating the initial encryption rule library in step S300 is as follows:
[0018] S310: Use the Abstract Syntax Tree (AST) parser to perform multi-level feature extraction on each source code file in the target code library to obtain encryption function feature groups. All encryption function feature groups together constitute the encryption function feature group set.
[0019] S320: Match each encryption function feature group in the encryption function feature group set with all current encryption algorithm features in the initial encryption rule base. If the class to which the encryption function feature group belongs exists in the current initial encryption rule base and the encryption function feature group contains an encryption algorithm, then the match is considered successful. Then, the successfully matched encryption function feature is added to the initial encryption rule base to obtain the updated initial encryption rule base; otherwise, it is discarded.
[0020] S330: After traversing all encryption function feature groups in the encryption function feature group set, the final updated initial encryption rule library is obtained. The final updated initial encryption rule library is then deduplicated and optimized to obtain the updated encryption rule library.
[0021] Preferably, the contents of the encryption function feature group obtained in S310 are as follows:
[0022] The Abstract Syntax Tree (AST) parser is used to parse the Import / ImportFrom nodes in each source code file to obtain the encrypted function reference relationships;
[0023] The Abstract Syntax Tree (AST) parser is used to parse the ClassDef node in each source code file to obtain the class declarations related to the encryption function.
[0024] The encryption function context stack is maintained through the FunctionDef in the encryption function: based on the obtained encryption function reference relationship and the encryption function related class declaration, when entering the FunctionDef node, the function name, start line number and end line number of the encryption function are pushed into the stack structure to which the encryption function belongs, and the context pointer of the current encryption function is updated; when exiting the FunctionDef node, the top element of the encryption function is popped, and the context of the encryption function is restored to the parent function of the encryption function.
[0025] The Abstract Syntax Tree (AST) parser is used to parse the Call node in each source code file to obtain the encrypted function call chain.
[0026] Each source code file contains an encryption function tag, encryption function reference relationships, encryption function related class declarations, encryption function context stack, and encryption function call chain, which together form an encryption function feature group.
[0027] Preferably, the encryption rule update library obtained in S300 further includes:
[0028] A feature merging algorithm is used to uniformly identify encryption algorithm features across source code files. For encryption algorithm features belonging to the same category in different source code files, a union operation of encryption algorithm features is performed.
[0029] Compared with the prior art, the present invention has at least the following advantages:
[0030] This invention achieves groundbreaking innovations in three aspects: rule dynamism, context awareness, and engineering efficiency by constructing a self-iterative rule base and a context-aware analysis engine.
[0031] 1. This invention effectively addresses the core shortcomings of traditional encryption detection technologies—namely, the rigidity of their rule bases and their inability to adapt to the rapid iteration of cryptography—by constructing a real-time feature extraction and dynamic rule base update mechanism based on Abstract Syntax Trees (ASTs). Specifically, when scanning a target codebase, the system automatically extracts encryption function feature groups through multi-level AST parsing (including Import nodes, ClassDef nodes, FunctionDef nodes, and Call nodes), and continuously updates the encryption rule base using incremental rule discovery and deduplication optimization strategies. This method can automatically capture the features of emerging encryption libraries without manual intervention, improving adaptability to the ever-evolving cryptographic ecosystem, overcoming the high false negative rate caused by the lag in updating fixed rule bases, and achieving self-evolution of detection capabilities.
[0032] 2. To address the high false positive rate and inaccurate business risk identification issues caused by the lack of context awareness in existing technologies, this invention innovatively introduces full-stack context modeling and cross-file tracing technology. During AST parsing, the system dynamically maintains the encryption function context stack to accurately record function nesting levels and scopes; it uses a class inheritance resolution algorithm to generate unique fully qualified name identifiers; and by constructing an encryption function relationship graph, it can associate encryption function calls and context information across files. This significantly reduces false negatives caused by incomplete context information and enables accurate tracing and risk assessment of the actual use scenarios of encryption algorithms in complex business logic.
[0033] 3. At the engineering practice level, the distributed processing architecture of this invention completely reshapes the performance standards for large-scale codebase detection. By integrating ZIP virtual file system direct reading technology with an automated source code filtering pipeline, the preprocessing stage avoids the overhead of entity decompression and effectively eliminates non-target files; a distributed AST parsing strategy is used to process feature extraction and rule update tasks in parallel. The output end uses JSONL streaming generation technology to reduce peak memory requirements. This architecture design enables the system to efficiently process codebases of tens of thousands or even larger scales, significantly reducing the detection time that traditional solutions cannot handle and significantly reducing memory consumption. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of the encryption rule update process of the present invention.
[0035] Figure 2 This is a code representation of the graph generation algorithm called in Python code form, which is an application of the present invention.
[0036] Figure 3 This is a flowchart illustrating the method of the present invention. Detailed Implementation
[0037] The present invention will now be described in further detail.
[0038] This invention obtains method call information and file reference information from the nodes of the abstract syntax tree. Based on this information, it obtains the call relationships between methods, thereby iteratively updating the rules, functions, and usage scenarios of traditional encryption algorithms. Specifically, in Python code, it addresses the problem of automated identification and risk assessment of traditional encryption algorithms in large-scale Python code libraries by constructing an iterative rule base and a context-sensitive syntax tree analysis engine, thereby achieving high-precision positioning and compliance verification of encryption usage scenarios.
[0039] See Figures 1-3 An adaptive encryption algorithm dynamic detection method based on abstract syntax trees includes the following steps:
[0040] S100: Construct an initial encryption rule base. Extract several encryption algorithm feature rules from publicly available cryptographic literature using a regularization matching algorithm to form the initial encryption rule base. The encryption algorithm features include encryption function labels and the class to which the encryption function belongs; the regularization matching algorithm is existing technology; general encryption algorithm features include the module import path list imports and its related features, the class declaration name list classes and its related features, the function signature list functions and its related features, etc.
[0041] S200: Select a public code repository, which includes source code files or a compressed code repository; preprocess the code repository to obtain a target code repository, where each source code file in the target code repository has an encrypted function tag;
[0042] The preprocessing of the code library in step S200 yields the following content for the target code library:
[0043] First, the source code file is checked; if it is the target source code file, it is kept; otherwise, it is filtered and discarded.
[0044] Secondly, the compressed code repository is checked. First, the directory structure of the code repository is decompressed to a temporary directory. Then, the source files in the temporary directory are checked to see if they are the target source files. If they are, they are kept; otherwise, they are filtered and discarded.
[0045] Finally, the retained target source code files are used to form the target code library.
[0046] The core function of the preprocessing step is to efficiently filter and standardize input. Through file type detection and decompression, a large number of non-source code files (such as binary files, documentation, and build artifacts) irrelevant to the analysis are effectively filtered out. This significantly reduces the amount of invalid data that needs to be processed in the subsequent abstract syntax tree (AST) scanning and analysis stages, avoiding parsing failures or introducing interference from irrelevant files. Simultaneously, it ensures that the input is plain text source code files in a specific programming language that can be directly parsed, providing a uniformly formatted and clean analysis object for subsequent rule updates, AST scanning, and graph construction. Finally, the compressed repository is processed while preserving the directory structure, maintaining the original modular structure and file reference relationships of the code project. This is crucial for accurately constructing a cross-file encrypted function call relationship graph.
[0047] S300: Dynamically update the initial encryption rule base to obtain an updated encryption rule base;
[0048] The encryption rule update library obtained in S300 also includes:
[0049] A feature merging algorithm is employed to uniformly identify encryption algorithm features across different source code files. For encryption algorithm features belonging to the same category in different source code files, a union operation is performed on the encryption algorithm features. A hash index mechanism is introduced to improve matching efficiency. In addition, a rule version control mechanism is introduced, automatically generating a snapshot each time a rule is updated, and supporting serialization and storage in JSON format for easy version tracking and rollback operations.
[0050] The contents of the updated encryption rule library obtained by dynamically updating the initial encryption rule library in S300 are as follows:
[0051] S310: The Abstract Syntax Tree (AST) parser is used to perform multi-level feature extraction on each source code file in the target code library to obtain encryption function feature groups. All encryption function feature groups together constitute the encryption function feature group set. The Abstract Syntax Tree (AST) parser is an existing technology.
[0052] The contents of the encryption function feature group obtained in S310 are as follows:
[0053] The Abstract Syntax Tree (AST) parser is used to parse the Import / ImportFrom nodes in each source code file to obtain the cryptographic function reference relationships. The cryptographic function reference relationships also contain the call location information of the cryptographic functions. This function not only extracts the reference relationships related to cryptographic functions, but also excludes non-cryptographic module information.
[0054] The Abstract Syntax Tree (AST) parser is used to parse the ClassDef node in each source code file to obtain the class declarations related to the encrypted functions. Parsing the ClassDef node to extract the encrypted class declarations involves using an attribute recursive parsing algorithm to generate a fully qualified name call chain identifier in the format of module.Class.function, recording metadata such as the start and end line numbers of the class definition.
[0055] The encryption function context stack is maintained through the FunctionDef in the encryption function: based on the obtained encryption function reference relationship and the encryption function related class declaration, when entering the FunctionDef node, the function name, start line number and end line number of the encryption function are pushed into the stack structure to which the encryption function belongs, and the context pointer of the current encryption function is updated; when exiting the FunctionDef node, the top element of the encryption function is popped, and the context of the encryption function is restored to the parent function of the encryption function.
[0056] The Abstract Syntax Tree (AST) parser is used to parse the Call node in each source code file to obtain the encrypted function call chain. For the Call node, the current encrypted function name, the line number of the call position, and its parent function and other related call chain information are parsed to construct an accurate context call relationship.
[0057] Each source code file contains an encryption function tag, encryption function reference relationships, encryption function related class declarations, encryption function context stack, and encryption function call chain, which together form an encryption function feature group.
[0058] This step employs a multi-layered AST parsing and context management mechanism to accurately extract encrypted function features and construct a structured feature set. Specifically: It parses Import / ImportFrom nodes to filter encryption-related references, clarify the location information of encrypted functions, and eliminate interference from non-encrypted modules; it parses ClassDef nodes using an attribute recursive algorithm to generate fully qualified call chain identifiers, capturing the definition location and metadata of encrypted class declarations; it dynamically maintains the encrypted function context stack, recording function nesting levels and scope to ensure the contextual coherence of the call chain; and it parses Call nodes to associate the current encrypted function with its parent function, constructing cross-level call relationships. These four parsing mechanisms work together to systematically extract a complete feature set containing encrypted function tags, encrypted function references, class declarations, context stacks, and call chains, providing the target codebase with non-redundant, structured encrypted behavior feature data, ensuring the accuracy and contextual integrity of encryption algorithm detection.
[0059] S320: Match each encryption function feature group in the encryption function feature group set with all current encryption algorithm features in the initial encryption rule base. If the class to which the encryption function feature group belongs exists in the current initial encryption rule base, and the encryption function feature group contains an encryption algorithm, then the match is considered successful. The successfully matched encryption function feature is then added to the initial encryption rule base to obtain the updated initial encryption rule base; otherwise, it is discarded. The matching method is to check whether the class to which the encryption function belongs is a class already included in the initial encryption rule base, and at the same time check whether the encryption function contains an encryption algorithm. Only after both of these conditions are met will it be added to the initial encryption rule base.
[0060] S330: After traversing all encryption function feature groups in the encryption function feature group set, the final updated initial encryption rule library is obtained. The final updated initial encryption rule library is then deduplicated and optimized to obtain the updated encryption rule library.
[0061] The core function of this traversal and deduplication optimization operation is to refine and improve the quality of the encryption rule base. By traversing all encryption function feature groups to ensure the complete inclusion of new rules, a deduplication operation is performed to effectively eliminate redundant entries in the rule base. This reduces the computational burden caused by repeated rule matching in subsequent scanning stages, improving detection efficiency. Simultaneously, it avoids the risk of duplicate counting or false alarms in detection results, ensuring the accuracy and reliability of the final output graph of encryption algorithm usage locations and their call relationships. This also lays an efficient and clear foundation for the continuous dynamic updating mechanism of the encryption rule base.
[0062] S400: Based on the encryption rule update library, the target code library is scanned again to locate the positions of all source code that uses encryption algorithms, obtaining a set of positions. An encryption function relationship graph is constructed, consisting of all encryption functions called by each source code. The relationship graphs of all source code encryption functions form a set. The code positions using encryption algorithms include the file path and line number of the source code. The encryption function relationship graph includes the parent function of the used encryption function and the class environment information of that encryption function. When constructing the relationship graph, the risks of using encryption algorithms are assessed based on dimensions such as whether they are called within security-sensitive functions, whether they are used in class initialization methods, and whether key hard-coding is involved. Then, a graph generation algorithm based on source code is called to construct the call relationship graph between encryption functions.
[0063] S500: Outputs the location set and relationship graph set to obtain the detection results of all source code that uses encryption algorithms. The results are output in JSONL streaming format, including algorithm name, risk level and complete call chain information.
[0064] Experimental content and results
[0065] Experimental content
[0066] 1. Comprehensive coverage and detection of traditional encryption algorithms
[0067] In software development, the use of traditional encryption algorithms is often hidden within complex code structures, making them difficult to identify directly. This invention innovatively collects and organizes relevant information on 17 traditional encryption algorithms through literature review, including their import statements, class definitions, and function calls in the Python language. For example, in Python, for the DES encryption algorithm, this invention analyzes in detail its common import methods such as Crypto.Cipher.DES, class definitions such as DES and DesKey, and function calls such as DES.new. This comprehensive collection of algorithm information lays a solid foundation for subsequent code detection, ensuring the comprehensiveness and accuracy of the detection.
[0068] 2. Application of Abstract Syntax Tree (AST) in Encryption Algorithm Detection
[0069] To address the problem that existing detection tools struggle to effectively identify encryption algorithm calls in code, this invention proposes an adaptive dynamic detection method for encryption algorithms based on Abstract Syntax Trees (ASTs). By developing an AST, this method can parse Python import statements and class definitions, as well as C header files, to accurately extract function calls related to 17 traditional encryption algorithms and update existing function information. This approach not only improves detection accuracy but also adapts to the characteristics of different programming languages, demonstrating powerful cross-language detection capabilities.
[0070] 3. Detection and scenario analysis of encryption algorithms in complex project libraries
[0071] In large project libraries, the calls to encryption algorithms are often hidden within complex file structures and code logic, making them difficult to detect effectively. This invention addresses this by using an Abstract Syntax Tree (AST) to traverse the project library under investigation. By analyzing import statements, class definitions, and function calls in the code, it accurately identifies the usage of encryption algorithms. Furthermore, AST analysis can trace the call paths of encryption algorithms within the code, revealing more specific scenarios and call relationships where the algorithm is used. This deep analysis identifies the use of encryption algorithms, reveals their actual application scenarios in the project, and provides crucial information for code optimization and security assessment.
[0072] 4. Efficiency and scalability of the detection method
[0073] The encryption algorithm detection method of this invention has significant advantages in terms of efficiency and scalability. Through an optimized AST parsing algorithm, it can achieve rapid analysis of large-scale codebases while ensuring detection accuracy. Experiments show that the core algorithm of this invention has low time complexity and can effectively handle the detection needs of large projects or complex codebases. Furthermore, this method also has good scalability, easily adapting to the addition of new encryption algorithms and programming languages, providing strong support for future code detection tasks.
[0074] Experimental evaluation
[0075] To scientifically evaluate the technical effectiveness of this invention, a rigorous experimental system was constructed on 15,000 real Python codebases. The experiment combined manual verification with automated analysis, with cryptography experts performing full annotation on 300 randomly sampled items to construct a benchmark dataset. The evaluation metrics adopted the gold standard in the field of information security detection: Precision measures the reliability of the detection results, Recall assesses the completeness of vulnerability coverage, and the F1-score comprehensively characterizes overall accuracy. Precision is based on the formula... Recall is calculated based on the formula The F1 score is calculated as follows: TP (True Positives) is the set of methods correctly identified as containing encryption algorithms; FP (False Positives) is the set of methods incorrectly identified as containing encryption algorithms; and FN (False Negatives) is the set of methods that cannot be identified as containing encryption algorithms. The F1 score is based on the formula... Pre and Re are abbreviations for Precision and Recall, respectively, representing the trade-off between the correctness and completeness of the results.
[0076] As shown in Table 1, this invention achieves a comprehensive breakthrough in the core performance indicators of encryption algorithm detection: the regularization extraction method has a recall rate as low as 45.0% and an F1-score of only 54.5% because it cannot cope with dynamic language characteristics and code structure variations; while this invention improves the accuracy to 96.3%, the recall rate to 89.0%, and the F1-score to 94.2% through dynamic rule evolution and context-aware engine.
[0077] Table 1. Performance comparison between the regularization extraction method and the method of this invention.
[0078] method Accuracy (%) Recall rate (%) F1-Score Regularized extraction 75.0% 45.0% 54.5% Method of the present invention 96.3% 89.0% 94.2%
[0079] In terms of time cost and engineering efficiency, this invention adopts a virtual file system solution, achieving a technological breakthrough in streaming ZIP processing. As shown in Table 2, the traditional "complete decompression and processing" solution requires 42.5 hours to complete a scan of 15,000 libraries, with a peak memory usage of 24.6GB and an interruption rate as high as 9.2%; while this invention reduces the total time to 6.8 hours (6.2 times efficiency improvement) and reduces the peak memory usage to 8.5GB (65.4% reduction in resource consumption).
[0080] Table 2 compares the engineering performance of traditional decompression scanning methods and the method of this invention.
[0081] System Solution Total time (hours) Peak memory usage (GB) Interruption rate (%) Decompress and process all files. 42.5 24.6 9.2 Method of the present invention 6.8 8.5 0
[0082] In summary, this invention proposes an adaptive encryption algorithm dynamic detection method based on abstract syntax trees. It solves the core problem of the disconnect between cryptographic evolution and detection rules by constructing a rule self-iteration mechanism, and combines a context-aware engine to solve the problem of call chain breakage caused by the dynamic language characteristics of Python. It also innovatively achieves accurate positioning in cross-file encryption scenarios. At the same time, relying on a virtual file system and streaming processing architecture, it maintains ultra-high efficiency even at the scale of tens of thousands of databases.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A dynamic detection method for adaptive encryption algorithms based on abstract syntax trees, characterized in that: Includes the following steps: S100: Construct an initial encryption rule base. Extract several encryption algorithm feature rules from publicly available cryptography literature using a regularization matching algorithm to form the initial encryption rule base. The encryption algorithm features include encryption function labels and the class to which the encryption function belongs. S200: Select a public code repository, which includes source code files or a code repository in compressed format; The codebase is preprocessed to obtain the target codebase, and each source code file in the target codebase has an encrypted function tag; S300: Dynamically update the initial encryption rule base to obtain an updated encryption rule base; S400: Based on the encryption rule update library, the target code library is scanned again to locate the positions of all code source code in the target code library that uses encryption algorithms, and a set of positions is obtained. An encryption function relationship graph is constructed, which consists of all encryption functions called by each code source code. The encryption function relationship graphs of all code source code form a relationship graph set. S500: Outputs the location set and relationship graph set to obtain the detection results for all source code that uses encryption algorithms.
2. The adaptive encryption algorithm dynamic detection method based on abstract syntax tree as described in claim 1, characterized in that: The preprocessing of the code library in step S200 yields the following content for the target code library: First, the source code file is checked; if it is the target source code file, it is kept; otherwise, it is filtered and discarded. Secondly, the compressed code repository is checked. First, the directory structure of the code repository is decompressed to a temporary directory. Then, the source files in the temporary directory are checked to see if they are the target source files. If they are, they are kept; otherwise, they are filtered and discarded. Finally, the retained target source code files are used to form the target code library.
3. The adaptive encryption algorithm dynamic detection method based on abstract syntax tree as described in claim 2, characterized in that: The contents of the updated encryption rule library obtained by dynamically updating the initial encryption rule library in S300 are as follows: S310: Use the Abstract Syntax Tree (AST) parser to perform multi-level feature extraction on each source code file in the target code library to obtain encryption function feature groups. All encryption function feature groups together constitute the encryption function feature group set. S320: Match each encryption function feature group in the encryption function feature group set with all current encryption algorithm features in the initial encryption rule base. If the class to which the encryption function feature group belongs exists in the current initial encryption rule base and the encryption function feature group contains an encryption algorithm, then the match is considered successful. Then, the successfully matched encryption function feature is added to the initial encryption rule base to obtain the updated initial encryption rule base; otherwise, it is discarded. S330: After traversing all encryption function feature groups in the encryption function feature group set, the final updated initial encryption rule library is obtained. The final updated initial encryption rule library is then deduplicated and optimized to obtain the updated encryption rule library.
4. The adaptive encryption algorithm dynamic detection method based on abstract syntax tree as described in claim 3, characterized in that: The contents of the encryption function feature group obtained in S310 are as follows: The Abstract Syntax Tree (AST) parser is used to parse the Import / ImportFrom nodes in each source code file to obtain the encrypted function reference relationships; The Abstract Syntax Tree (AST) parser is used to parse the ClassDef node in each source code file to obtain the class declarations related to the encryption function. The context stack of the encryption function is maintained through FunctionDef in the encryption function: by obtaining the encryption function reference relationship and the class declaration related to the encryption function, when entering the FunctionDef node, the function name, start line number and end line number of the encryption function are pushed into the stack structure to which the encryption function belongs, and the context pointer of the current encryption function is updated. When exiting the FunctionDef node, pop the top element of the encryption function's stack and restore the context of the encryption function's parent function. The Abstract Syntax Tree (AST) parser is used to parse the Call node in each source code file to obtain the encrypted function call chain. Each source code file contains an encryption function tag, encryption function reference relationships, encryption function related class declarations, encryption function context stack, and encryption function call chain, which together form an encryption function feature group.
5. The adaptive encryption algorithm dynamic detection method based on abstract syntax trees as described in claim 4, characterized in that: The encryption rule update library obtained in S300 also includes: A feature merging algorithm is used to uniformly identify encryption algorithm features across source code files. For encryption algorithm features belonging to the same category in different source code files, a union operation of encryption algorithm features is performed.
Citation Information
Patent Citations
Tree structure-based cryptographic algorithm logical expression identification method
CN102799806A
Code rule checking method and system based on knowledge base feature matching
CN113849413A
Code cryptographic algorithm type identification and parameter misuse detection method and system
CN115658542A
Code analysis tool for recommending encryption of data without affecting program semantics
US20160254911A1
Evaluating Cryptographic API Calls at Runtime
US20250227123A1