Embedded code vulnerability detection method and device, computer equipment and medium
By employing a multi-engine collaborative verification and adaptive learning framework, the problems of high false positive rate, high resource consumption, and insufficient path coverage in embedded code vulnerability detection are solved, achieving efficient vulnerability detection and remediation with low memory usage.
Patent Information
- Application Number
- CN202510911070.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for embedded code vulnerability detection suffer from problems such as high false positive rates, incomplete rule coverage, high resource consumption, insufficient path coverage, and strong data dependencies.
A multi-engine collaborative verification mechanism is adopted, which combines a rule engine, an AI engine, and a symbolic execution engine for weighted calculation to generate weighted detection results. False alarms are filtered through a central control module, path pruning is performed using a lightweight symbolic execution engine, and weights are dynamically adjusted using an adaptive learning framework to achieve vulnerability remediation.
It reduces the false positive rate of vulnerability detection, reduces memory usage, improves path coverage, is suitable for embedded devices, shortens the specification support cycle, and improves detection efficiency and remediation speed.
Smart Images

Figure CN120951331A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of code detection technology, and in particular to a method, apparatus, computer device, and medium for detecting embedded code vulnerabilities. Background Technology
[0002] Currently, the field of static vulnerability detection for embedded code (such as C / C++) mainly adopts the following three types of technical solutions: 1. Static analysis based on rule matching.
[0003] Technical principle: Code pattern matching is performed using predefined coding standards (such as MISRA C / C++, CERT C) and vulnerability pattern libraries, combined with lexical analysis, data flow analysis and other technologies.
[0004] Typical tools: Fortify SCA: Supports 4100+ diagnostic rules, covering standards such as MISRA C:2012, but rule updates lag behind newer specifications (such as AUTOSAR C++14).
[0005] Cppcheck: Detects memory leaks and buffer overflows through data stream analysis, with a false positive rate as high as 42%.
[0006] Implementation process: Syntax tree construction (Clang AST), control flow graph (CFG) generation, static rule matching (such as verification of strcpy function usage).
[0007] 2. Symbolic execution and dynamic analysis.
[0008] Technical principle: Verify program behavior by symbolizing the execution path or by monitoring runtime.
[0009] Symbolic execution: Tools such as KLEE generate path constraints and detect buffer overflows, but there is a path explosion problem (it fails when the number of paths is >10^5).
[0010] Dynamic analysis: Valgrind: Monitors runtime memory errors, but increases memory usage by 40%-60%, making it unsuitable for embedded real-time systems.
[0011] Fuzzing: AFL triggers anomalies through input mutation, but its coverage of embedded binary files is insufficient (only 30%-50% of the paths are covered).
[0012] 3. Machine learning-assisted detection Technical principle: Use historical vulnerability data (such as CVE database) to train a model to identify abnormal patterns.
[0013] Pattern recognition: Coverity: Combining statistical models and static analysis, the false positive rate is reduced to 15%, but it relies on high-quality labeled data (CVE vulnerability database coverage <30%).
[0014] TscanCode identifies buffer overflows based on a bidirectional long short-term memory (LSTM) network model, but has poor generalization ability for unseen vulnerability types.
[0015] The drawbacks of rule-based static analysis include a high false positive rate (traditional tools have a false positive rate >35%, such as Cppcheck at 42%) and incomplete rule coverage (mainstream tools support an average of 2000+ rules, lagging behind in support of newer specifications such as AUTOSAR C++14). Symbolic execution and dynamic analysis suffer from high resource consumption (dynamic analysis increases memory usage by 40%-60%, failing to meet the resource limitations of embedded devices) and insufficient path coverage (fuzz testing has a <50% coverage rate for embedded binary paths). Machine learning-assisted detection is highly dependent on data (requires annotation of historical vulnerability data (CVE library coverage <30%), and has poor generalization ability for new vulnerabilities) and lacks remediation guidance (90% of tools only provide general suggestions and lack automatic patch generation capabilities). Summary of the Invention
[0016] In view of this, embodiments of the present invention provide a method for detecting embedded code vulnerabilities to address the technical problems in existing technologies, such as high false positive rates and incomplete rule coverage in static analysis based on rule matching, high resource consumption and insufficient path coverage in symbolic execution and dynamic analysis, and strong data dependence and lack of remediation guidance in machine learning-assisted detection. The method includes: The embedded code to be detected is input into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module performs weighted calculations on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. The weighted detection results and the project feature library are input into the central control module. Based on the severity and impact of the vulnerabilities in the weighted detection results, the detected vulnerabilities are sorted and false positives are filtered to generate a list of vulnerabilities to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. The interactive repair system repairs vulnerabilities in the vulnerability list and replaces insecure function calls in the embedded code, generating code repair results.
[0017] This invention also provides an embedded code vulnerability detection device to address the technical problems in existing technologies, such as high false positive rates and incomplete rule coverage in static analysis based on rule matching, high resource consumption and insufficient path coverage in symbolic execution and dynamic analysis, and strong data dependence and lack of remediation guidance in machine learning-assisted detection. The device includes: The vulnerability detection module is used to input the embedded code to be detected into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module performs weighted calculation on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. The vulnerability list generation module is used to input the weighted detection results and the project feature library into the central control module. Based on the severity and impact of the vulnerabilities in the weighted detection results, the detected vulnerabilities are sorted and false positives are filtered to generate a vulnerability list to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. The vulnerability repair module is used to repair vulnerabilities in the vulnerability list through an interactive repair system, replace insecure function calls in the embedded code, and generate code repair results.
[0018] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for detecting any of the embedded code vulnerabilities, thereby solving the technical problems in the prior art, such as high false alarm rate and incomplete rule coverage of static analysis based on rule matching, high resource consumption and insufficient path coverage of symbolic execution and dynamic analysis, and strong data dependence and lack of remediation guidance of machine learning-assisted detection.
[0019] This invention also provides a computer-readable storage medium storing a computer program that executes any of the above-described methods for detecting embedded code vulnerabilities, in order to solve the technical problems in the prior art, such as high false alarm rate and incomplete rule coverage in static analysis based on rule matching, high resource consumption and insufficient path coverage in symbolic execution and dynamic analysis, and strong data dependence and lack of remediation guidance in machine learning-assisted detection.
[0020] Compared with the prior art, the beneficial effects that at least one technical solution adopted in the embodiments of this specification can achieve include at least: The false positive rate of vulnerability detection is reduced through a multi-engine detection and collaborative verification mechanism; memory usage is reduced through a lightweight symbolic execution engine (taint analysis-driven path pruning), making it suitable for embedded devices. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an embedded code vulnerability detection method provided by an embodiment of the present invention; Figure 2 This is a system architecture diagram of an embodiment of the embedded code vulnerability detection method provided by the present invention; Figure 3 This is a flowchart of the multi-engine collaborative analysis module provided in an embodiment of the present invention; Figure 4 This is a flowchart of the adaptive learning framework provided in an embodiment of the present invention; Figure 5 This is a flowchart of the central control module provided in an embodiment of the present invention; Figure 6 This is a flowchart of the interactive repair system provided in the embodiments of the present invention; Figure 7 This is a structural block diagram of a computer device provided in an embodiment of the present invention; Figure 8 This is a structural block diagram of an embedded code vulnerability detection device provided in an embodiment of the present invention. Detailed Implementation
[0023] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0024] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] In this embodiment of the invention, a method for detecting embedded code vulnerabilities is provided, such as... Figure 1 and Figure 2As shown, the method includes: Step S101: Input the embedded code to be detected into the multi-engine analysis synchronization module to generate multi-engine detection results. The central control module performs weighted calculation on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. Step S102: Input the weighted detection results and project feature library into the central control module. Based on the severity and impact of the vulnerabilities in the weighted detection results, sort the detected vulnerabilities and filter the false positive detection results to generate a list of vulnerabilities to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. Step S103: Repair the vulnerabilities in the vulnerability list through the interactive repair system, replace the insecure function calls in the embedded code, and generate code repair results.
[0026] In specific implementation, the embedded code to be detected is input into the multi-engine analysis and synchronization module through the following steps to generate multi-engine detection results. The central control module then performs a weighted calculation on the multi-engine detection results to generate a weighted detection result: The embedded code is subjected to rule verification using a rule engine based on a pointer alias analysis algorithm, generating static analysis results. Rule confidence is calculated from these static analysis results. The embedded code is then subjected to vulnerability detection and identification using an AI engine, generating vulnerability detection results. AI confidence is calculated from these vulnerability detection results. The embedded code is then subjected to detection using a lightweight symbolic execution engine, generating symbolic execution verification results. A rule confidence weighting rate, an AI confidence weighting rate, and a symbolic execution verification confidence weighting rate are set. A weighted detection result is calculated using these weighted rates, rule confidence, AI confidence weighting rate, AI confidence, symbolic execution verification confidence weighting rate, and the symbolic execution verification results.
[0027] The processing flow of the multi-engine analysis synchronization module is as follows: Figure 3 As shown.
[0028] Specifically, the rule engine is based on 5,000+ coding standards such as MISRA C:2023 and AUTOSAR C++14 (extended from 4,107 rules of Helix QAC), and performs static pattern matching through syntax trees (Clang AST) and control flow graphs (CFG).
[0029] By introducing a pointer aliasing analysis algorithm and using the dynamic pointer propagation graph (DPPG) model in the cross-function analysis technique of the three-level pointer tracking matrix, the problem of misjudgment of multi-level pointer operations in traditional static analysis is solved (such as the accuracy of out-of-bounds detection of char **p is improved by 40%). Compared with the traditional Stensgaard algorithm, the detection accuracy of multi-level pointer operations is improved from 62% to 92%.
[0030] Specifically, the AI detection engine uses a bidirectional long short-term memory (LSTM) network model (or a Transformer model) to analyze code sequences, and is trained in conjunction with a historical CVE vulnerability database to support the identification of vulnerabilities such as buffer overflows and null pointer dereferences (for example, the training data covers the CWE Top 25 2023 vulnerability types). It also incorporates a graph neural network (GNN) to construct a function call graph (FCG) embedding, enhancing its ability to detect multi-threaded race conditions.
[0031] The symbolic execution engine employs a lightweight design, utilizing a path pruning algorithm driven by taint analysis (which tracks the propagation path of external input data in the code to identify unverified sensitive operations, such as those proposed in the symbolic execution optimization) to compress the symbolic execution memory footprint to below 50MB (compared to 200MB+ for traditional tools like KLEE). Dynamic verification is employed, generating symbolic constraints for buffer overflow paths and combining fuzzy testing seeds to optimize coverage (achieving 80% path coverage).
[0032] The three-engine collaborative mechanism employs a confidence-weighted algorithm (e.g., Final_Score = 0.4 * rule confidence + 0.3 * AI confidence + 0.3 * symbolic execution verification result), reducing the false positive rate from 42% for traditional tools to 7.8%.
[0033] Parallel processing reduces detection time by 166% (from 30 minutes serially to 11 minutes in parallel). Specifically, three engines (rule engine, AI engine, and symbolic execution engine) execute in parallel, analyzing the same code segment simultaneously. Furthermore, through dynamic task distribution, the central control module breaks down the code into subtasks by function / code block. Resources are allocated based on complexity (e.g., high-complexity code blocks are prioritized for the symbolic execution engine). An event-driven process is implemented: after an engine completes a subtask, it sends an event → the central module summarizes the results → triggers a confidence-weighted calculation.
[0034] In practice, the following steps are used to update the project signature database to help increase the vulnerability detection rate and remediation effectiveness: The system acquires the characteristics of new vulnerabilities in the vulnerability list, extracts the characteristics of the new vulnerabilities, generates vulnerability pattern fingerprints, and updates the vulnerability pattern fingerprints to the historical vulnerability database. The vulnerability pattern fingerprints are used to identify and match similar vulnerability patterns. The system updates the historical vulnerability distribution heatmap based on the multi-engine detection results and updates the historical vulnerability distribution heatmap to the project feature database. The system dynamically adjusts the weights of the software architecture standards used for vulnerability detection using a dynamic weight adjuster.
[0035] In practice, the following steps are used to dynamically adjust the weights of the software architecture standards used for vulnerability detection via a dynamic weight adjuster: The code analysis of a set number of lines is used as a sliding window. The weights of the software architecture standards used for vulnerability detection are recalculated as an adjustment rule when the sliding window changes. An analysis mode is selected based on the code size of the embedded code, including either a deep analysis mode or a standard analysis mode. If the analysis mode is the deep analysis mode, the current vulnerability detection rate is calculated. Based on the adjustment rule, the weights of the software architecture standards used for vulnerability detection are dynamically adjusted using the sliding window mean method, according to the vulnerability detection rate, and subsequent vulnerability detection is performed based on the adjusted weights.
[0036] Specifically, the historical vulnerability database integrates the CVE vulnerability signature database (covering 300,000+ samples from the CWE Top 25 in 2023), and generates vulnerability pattern fingerprints (such as the strcpy unchecked length buffer overflow pattern) through feature extraction.
[0037] The project feature library quantifies project characteristics. It stores code cyclomatic complexity (a metric that measures the complexity of code logic branches and is used for project feature library modeling in proposals), module coupling, and historical vulnerability distribution heatmaps (visualized project-level vulnerability distribution maps generated based on historical vulnerability databases).
[0038] like Figure 4 As shown, code scanning uses dynamic adaptation. In one implementation, when the code size exceeds 100,000 lines, a deep analysis mode is automatically triggered (such as the MISRA rule weighting for automotive electronics code enhancement).
[0039] A dynamic weight adjuster is used to dynamically adjust the weights. In one implementation, the sliding window mean method is used to update the rule weights (window size = 1000 lines of code), and the rule priority is dynamically adjusted according to the vulnerability detection rate (e.g., AUTOSAR rule weight + 0.2).
[0040] Enhance model generalization ability by leveraging unlabeled code data to address the issue of insufficient CVE vulnerability database coverage (<30%). Evaluate remediation effectiveness using FeedbackAgent and dynamically update vulnerability distribution heatmaps in the project feature library.
[0041] In practice, the following steps are used to sort the detected vulnerabilities based on the severity and impact of the weighted detection results, filter out false positives, and generate a list of vulnerabilities to be fixed: Obtain the detection results of the rule engine and the AI engine; determine whether the detection results of the rule engine and the AI engine conflict; if a conflict exists, initiate symbolic execution verification, filter the detection results of the rule engine and the AI engine with confidence levels lower than a set confidence level based on a voting mechanism, and generate filtered detection results; adjust the weight of the AI engine's detection results according to the code cyclomatic complexity of the project feature library, and generate adjusted detection results; conduct a remediation impact assessment on the adjusted detection results, classify the filtered detection results into vulnerabilities according to the CWE severity level, sort the adjusted detection results according to the vulnerability classification results and the remediation impact assessment results, and generate a list of vulnerabilities to be remediated.
[0042] like Figure 5 As shown, false positive filtering based on a voting mechanism is used to filter low-confidence results (symbolic execution verification is triggered when the results of the rule engine and the AI engine conflict).
[0043] After classifying vulnerabilities, a priority ranking is generated. Specifically, a priority queue is generated based on the CWE severity level (Critical / High / Medium) and the remediation impact assessment results (remediation impact score) (e.g., ranking basis = 0.6 × CWE level + 0.4 × remediation impact score).
[0044] The weights are dynamically adjusted by combining the confidence-weighted algorithm with the code complexity of the project feature library (for example, the weight of the AI engine is reduced when the code cyclomatic complexity is >15).
[0045] It adopts a lightweight message queue, supports 10+ concurrent task processing, and reduces memory usage by 40%.
[0046] In practice, the following steps are used to repair vulnerabilities in the vulnerability list through an interactive repair system, replace insecure function calls in the embedded code, and generate code repair results: If the selected repair type is code patch, replace the function that caused the vulnerability with a secure function to generate a code patch, and use the code patch to repair the vulnerability; if the vulnerability needs to be repaired by adjusting the configuration, configure the compiler options and then restart the service to verify the repair; if the architecture needs to be optimized, repair the vulnerability by decoupling the embedded code into modules; use a repair impact assessment model based on random forest regression algorithm to perform regression testing on the system after the vulnerability is repaired to check whether the embedded code after the vulnerability is repaired can still run normally; If the test runs normally, the vulnerability features corresponding to the patched vulnerability will be updated to the project feature library. If the test fails, a backtracking report including the reason for the test failure will be generated.
[0047] like Figure 6 As shown, automatic replacement is achieved through a safe function replacer (e.g., strcpy → strncpy (buffer overflow repair), malloc → calloc (uninitialized memory repair)). The safe function replacer covers the safe function replacement rule library of the C99 / C++17 standard library and contains alternatives for 87 high-risk functions.
[0048] By using a repair impact assessment model based on the random forest regression algorithm (training data from 1000+ open source projects), the error in predicting memory consumption after repair is reduced to less than 5%.
[0049] Generate a comparison chart of memory usage before and after the repair, as well as a curve showing the change in execution cycle, as a visualization report (refer to the feedback mechanism of HybridAgent).
[0050] The interactive remediation system provides three levels of remediation suggestions: code patching, configuration adjustments (such as hardening compilation options), and architecture optimization (such as module decoupling). Clicking on a remediation suggestion will trace back to the original vulnerable line of code and associated rule entries (refer to the HelixQAC Dashboard feature).
[0051] A multi-engine collaborative analysis module implements a three-dimensional collaborative architecture consisting of a rule engine (5000+ extended rules), an AI detection engine (Bidirectional Long Short-Term Memory Network (LSTM) + GNN), and a lightweight symbolic execution engine. Combined with a confidence-weighted algorithm and an event-driven parallel acceleration mechanism, this effectively improves detection rate and efficiency. The engine collaborative verification mechanism reduces the false positive rate from 42% (Cppcheck) of traditional tools to 7.8%. The lightweight symbolic execution engine (taint analysis-driven path pruning) reduces memory usage to 50MB (KLEE requires 200MB+), making it suitable for embedded devices.
[0052] By using the dynamic weight adjustment system (dynamic weight adjuster) of the adaptive learning framework, the rule priority is dynamically adjusted based on the sliding window mean method (window = 1000 LOC). Combined with the adaptive weight strategy of the project feature fingerprint library (code cyclomatic complexity / historical vulnerability heatmap), the support cycle for new specifications such as AUTOSAR C++14 is shortened to 2 weeks (traditional tools require 3-6 months).
[0053] Static taint analysis-driven symbolic execution optimization, a coverage improvement technique that combines static taint analysis and symbolic execution (path coverage increased from 30-50% (AFL) to 80%). Utilizing a path pruning algorithm, symbolic execution is performed only on high-risk paths marked with taints, skipping irrelevant paths, reducing memory usage to 50MB (75% lower than traditional tools).
[0054] An interactive remediation system enables automatic replacement of secure functions (e.g., strcpy → strncpy) and a three-tiered remediation suggestion system. A random forest-based remediation impact prediction model (with an error rate of <5%) and a visual source tracing mechanism are implemented. Automatic secure function replacement and multimodal remediation suggestions reduce remediation time from an average of 4.2 hours per vulnerability (Coverity) to 1.5 hours.
[0055] By using pointer alias analysis enhancement techniques (embedding pointer tracing algorithms during the Clang AST parsing stage), the accuracy of out-of-bounds detection for multi-level pointer operations (such as char**p) is improved by 40%.
[0056] In this embodiment, a computer device is provided, such as... Figure 7 As shown, it includes a memory 701, a processor 702, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for detecting any of the aforementioned embedded code vulnerabilities.
[0057] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.
[0058] In this embodiment, a computer-readable storage medium is provided, which stores a computer program that executes any of the above-described methods for detecting embedded code vulnerabilities.
[0059] Specifically, computer-readable storage media include both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media do not include transient media, such as modulated data signals and carrier waves.
[0060] Based on the same inventive concept, this invention also provides an embedded code vulnerability detection device, as described in the following embodiments. Since the principle of the embedded code vulnerability detection device in solving the problem is similar to that of the embedded code vulnerability detection method, the implementation of the embedded code vulnerability detection device can refer to the implementation of the embedded code vulnerability detection method, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0061] Figure 8 This is a structural block diagram of an embedded code vulnerability detection device according to an embodiment of the present invention, such as... Figure 8 As shown, it includes: vulnerability detection module 801, vulnerability list generation module 802, and vulnerability repair module 803. The structure is described below.
[0062] The vulnerability detection module 801 is used to input the embedded code to be detected into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module performs weighted calculation on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. The vulnerability list generation module 802 is used to input the weighted detection results and the project feature library into the central control module, sort the detected vulnerabilities and filter the false positive detection results according to the severity and impact of the vulnerabilities in the weighted detection results, and generate a vulnerability list to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. The vulnerability repair module 803 is used to repair vulnerabilities in the vulnerability list through an interactive repair system, replace insecure function calls in the embedded code, and generate code repair results.
[0063] In one embodiment, the vulnerability detection module includes: The rule confidence calculation unit is used to perform rule verification on the embedded code through a rule engine based on a pointer alias analysis algorithm, generate static analysis results of the rule engine, and calculate the rule confidence based on the static analysis results of the rule engine. The AI confidence calculation unit is used to perform vulnerability detection and identification on the embedded code through the AI engine, generate the vulnerability detection results of the AI engine, and calculate the AI confidence score based on the vulnerability detection results of the AI engine. The symbolic execution verification unit is used to detect the embedded code through a lightweight symbolic execution engine and generate symbolic execution verification results; The confidence calculation unit is used to set the rule confidence weighting rate, the AI confidence weighting rate, and the symbolic execution verification confidence weighting rate, and to calculate the weighted detection result using the rule confidence weighting rate, the rule confidence, the AI confidence weighting rate, the AI confidence, the symbolic execution verification confidence weighting rate, and the symbolic execution verification result.
[0064] In one embodiment, the symbolic execution verification unit is configured to analyze the propagation path of marked external input data using a static taint algorithm, and divide the propagation path into high-risk paths and irrelevant paths; detect the embedded code on the high-risk path using a lightweight symbolic execution engine and fuzzing, and generate symbolic execution verification results, wherein the fuzzing includes: performing static taint analysis on the embedded code, generating taint analysis results, optimizing the fuzzing seed using the taint analysis results, and using the optimized fuzzing seed as input for fuzzing.
[0065] In one embodiment, the vulnerability list generation module includes: The detection result acquisition unit is used to acquire the detection results of the rule engine and the detection results of the AI engine; The detection result filtering unit is used to determine whether the detection results of the rule engine and the detection results of the AI engine conflict. If there is a conflict, symbolic execution verification is initiated, and the detection results of the rule engine and the AI engine with confidence levels lower than the set confidence level are filtered based on the voting mechanism to generate filtered detection results. The weight adjustment unit is used to adjust the weights of the detection results of the AI engine according to the code cyclomatic complexity of the project feature library, and generate the adjusted detection results. The vulnerability list sorting unit is used to assess the remediation impact of the adjusted detection results, classify the filtered detection results into vulnerabilities according to the CWE severity level, sort the adjusted detection results according to the vulnerability classification results and the remediation impact assessment results, and generate a vulnerability list to be remediated.
[0066] In one embodiment, the vulnerability remediation module includes: The code patch generation unit is used to replace the function that caused the vulnerability with a security function if the selected repair type is code patch, generate a code patch, and use the code patch to repair the vulnerability. The configuration adjustment unit is used to modify the compiler options and then restart the service to verify that the vulnerability has been fixed if it is necessary to fix the vulnerability by adjusting the configuration. The module decoupling unit is used to fix vulnerabilities by decoupling embedded code modules if architecture optimization is required. The regression testing unit is used to perform regression testing on the system after vulnerability patching using a patching impact assessment model based on the random forest regression algorithm, to detect whether the embedded code after the vulnerability patching can still run normally; The first feature library update unit is used to update the vulnerability features corresponding to the repaired vulnerabilities to the project feature library if it can run normally, and to generate a backtracking report including the reason for the test failure if the test fails.
[0067] In one embodiment, the above-described apparatus further includes an adaptive learning module.
[0068] In one embodiment, the adaptive learning module includes: The vulnerability database update unit is used to obtain the characteristics of new vulnerabilities in the vulnerability list, extract the characteristics of the new vulnerabilities, generate vulnerability pattern fingerprints, and update the vulnerability pattern fingerprints to the historical vulnerability database. The vulnerability pattern fingerprints are used to identify and match similar vulnerability patterns. The second feature library update unit is used to update the historical vulnerability distribution heatmap according to the multi-engine detection results, and update the historical vulnerability distribution heatmap to the project feature library; The dynamic weight adjustment unit is used to dynamically adjust the weights of the software architecture standards used for vulnerability detection through a dynamic weight adjuster.
[0069] In one embodiment, the confidence calculation unit is configured to use a set number of lines of code analyzed at one time as a sliding window, and recalculate the weights of the software architecture standards used for vulnerability detection as an adjustment rule when the sliding window changes; select an analysis mode based on the code size of the embedded code, wherein the analysis mode is a deep analysis mode or a standard analysis mode; if the analysis mode is the deep analysis mode, calculate the current vulnerability detection rate; based on the adjustment rule, use the sliding window mean method to dynamically adjust the weights of the software architecture standards used for vulnerability detection according to the vulnerability detection rate, and perform subsequent vulnerability detection based on the adjusted weights.
[0070] The embodiments of this invention achieve the following technical effects: reducing the false positive rate of vulnerability detection through an engine-based collaborative verification mechanism; reducing memory usage through a lightweight symbolic execution engine (taint analysis-driven path pruning), suitable for embedded devices; dynamically adjusting rule priorities based on the sliding window mean method through an adaptive learning framework's dynamic weight adjustment system, and reducing the specification support cycle by combining it with a project feature fingerprint library, effectively saving processing time; improving path coverage in vulnerability detection through a coverage enhancement technique that links static taint analysis with symbolic execution; and reducing memory usage by using a path pruning algorithm to perform symbolic execution only on high-risk paths marked with taints, skipping irrelevant paths.
[0071] Obviously, those skilled in the art should understand that the modules or steps of the above-described embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
[0072] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting embedded code vulnerabilities, characterized in that, include: The embedded code to be detected is input into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module performs weighted calculations on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. The weighted detection results and the project feature library are input into the central control module. Based on the severity and impact of the vulnerabilities in the weighted detection results, the detected vulnerabilities are sorted and false positives are filtered to generate a list of vulnerabilities to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. The interactive repair system repairs vulnerabilities in the vulnerability list and replaces insecure function calls in the embedded code, generating code repair results.
2. The embedded code vulnerability detection method as described in claim 1, characterized in that, The embedded code to be detected is input into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module then performs a weighted calculation on the multi-engine detection results to generate a weighted detection result, including: The embedded code is validated by a rule engine based on a pointer alias analysis algorithm, and static analysis results of the rule engine are generated. The rule confidence is then calculated based on the static analysis results of the rule engine. The embedded code is subjected to vulnerability detection and identification by an AI engine, and the vulnerability detection results of the AI engine are generated. The AI confidence level is calculated based on the vulnerability detection results of the AI engine. The embedded code is tested using a lightweight symbolic execution engine, and symbolic execution verification results are generated. Set a rule confidence weighting rate, an AI confidence weighting rate, and a symbolic execution verification confidence weighting rate. Calculate the weighted detection result using the rule confidence weighting rate, the rule confidence level, the AI confidence weighting rate, the AI confidence level, the symbolic execution verification confidence weighting rate, and the symbolic execution verification result.
3. The embedded code vulnerability detection method as described in claim 2, characterized in that, The embedded code is inspected using a lightweight symbolic execution engine, generating symbolic execution verification results, including: The propagation path of external input data is analyzed using a static taint algorithm, and the propagation path is divided into high-risk paths and irrelevant paths. The embedded code on the high-risk propagation path is detected by a lightweight symbolic execution engine and fuzzing, and symbolic execution verification results are generated. The fuzzing includes: performing static taint analysis on the embedded code, generating taint analysis results, optimizing the fuzzing seed based on the taint analysis results, and using the optimized fuzzing seed as the input for fuzzing.
4. The embedded code vulnerability detection method as described in claim 1, characterized in that, Based on the severity and impact of the vulnerabilities in the weighted detection results, the detected vulnerabilities are sorted and false positives are filtered out to generate a list of vulnerabilities to be patched, including: Obtain the detection results from the rule engine and the detection results from the AI engine; Determine whether the detection results of the rule engine and the detection results of the AI engine conflict. If there is a conflict, start symbolic execution verification, filter the detection results of the rule engine and the AI engine with confidence scores lower than the set confidence scores based on the voting mechanism, and generate filtered detection results. The weights of the AI engine's detection results are adjusted based on the code cyclomatic complexity of the project feature library to generate adjusted detection results; An impact assessment of the remediation effects is performed on the adjusted detection results. The filtered detection results are classified into vulnerability levels according to their CWE severity. The adjusted detection results are then sorted based on the vulnerability classification results and the impact assessment results to generate a list of vulnerabilities to be remediated.
5. The embedded code vulnerability detection method as described in claim 1, characterized in that, The interactive repair system repairs vulnerabilities in the vulnerability list and replaces insecure function calls in the embedded code, generating code repair results, including: If the selected repair type is code patch, replace the function that caused the vulnerability with a secure function to generate a code patch, and use the code patch to repair the vulnerability; If you need to fix the vulnerability by adjusting the configuration, configure the compiler options and then restart the service to verify that the vulnerability has been fixed. If the architecture needs to be optimized, vulnerabilities can be fixed by decoupling the embedded code into modules. The system after vulnerability patching is subjected to regression testing using a patching impact assessment model based on random forest regression algorithm to detect whether the embedded code after the vulnerability patching can still run normally. If the test runs normally, the vulnerability features corresponding to the patched vulnerability will be updated to the project feature library. If the test fails, a backtracking report including the reason for the test failure will be generated.
6. The method for detecting embedded code vulnerabilities as described in any one of claims 1 to 5, characterized in that, Also includes: The features of new vulnerabilities in the vulnerability list are obtained, the features of the new vulnerabilities are extracted, a vulnerability pattern fingerprint is generated, and the vulnerability pattern fingerprint is updated to the historical vulnerability database. The vulnerability pattern fingerprint is used to identify and match similar vulnerability patterns. The historical vulnerability distribution heatmap is updated based on the multi-engine detection results, and the historical vulnerability distribution heatmap is updated to the project feature library. The weights of the software architecture standards used for vulnerability detection are dynamically adjusted using a dynamic weight adjuster.
7. The embedded code vulnerability detection method as described in claim 6, characterized in that, The weights of the software architecture standards used for vulnerability detection are dynamically adjusted by a dynamic weight adjuster, including: The code of a set number of lines is analyzed at one time as a sliding window, and the weight of the software architecture standard used for vulnerability detection is recalculated as an adjustment rule when the sliding window changes. The analysis mode is selected based on the code size of the embedded code, wherein the analysis mode is either a deep analysis mode or a standard analysis mode; If the analysis mode is the deep analysis mode, calculate the current vulnerability detection rate; Based on the aforementioned adjustment rules, the sliding window mean method is used to dynamically adjust the weights of the software architecture standards used for vulnerability detection according to the vulnerability detection rate, and subsequent vulnerability detection is performed based on the adjusted weights.
8. An embedded code vulnerability detection device, characterized in that, include: The vulnerability detection module is used to input the embedded code to be detected into the multi-engine analysis and synchronization module to generate multi-engine detection results. The central control module performs weighted calculation on the multi-engine detection results to generate weighted detection results. The multi-engine detection results include the static analysis results of the rule engine, the vulnerability detection results of the AI engine, and the verification results of the symbolic execution engine. The vulnerability list generation module is used to input the weighted detection results and the project feature library into the central control module. Based on the severity and impact of the vulnerabilities in the weighted detection results, the detected vulnerabilities are sorted and false positives are filtered to generate a vulnerability list to be fixed. The project feature library is used to store code cyclomatic complexity and historical vulnerability distribution heatmaps. The vulnerability repair module is used to repair vulnerabilities in the vulnerability list through an interactive repair system, replace insecure function calls in the embedded code, and generate code repair results.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the embedded code vulnerability detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the embedded code vulnerability detection method according to any one of claims 1 to 7.