Software security vulnerability detection method and device, equipment and storage medium
By automatically parsing the dependency declaration files of low-code projects and combining them with a dynamic security knowledge base and a large language model, a risk assessment report is generated, which solves the problem of high false positive rates in low-code projects and achieves efficient and accurate security detection for low-code development platforms.
Patent Information
- Application Number
- CN202511609348.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-03-06
AI Technical Summary
Low-code projects have a high false alarm rate in security detection. Existing tools lack risk analysis of real-world scenarios for low-code projects, resulting in a high false alarm rate and making it difficult to meet the needs of accurate security management.
By automatically triggering security checks, parsing dependency declaration files to obtain a list of dependencies to be checked, and combining a dynamic security knowledge base and a large language model, a risk detection and assessment report is generated, enabling scenario-based and precise risk analysis.
It reduces the false alarm rate of security detection, improves the automation, accuracy and intelligence of detection, ensures the timeliness and accuracy of vulnerability information, and supports efficient security management of low-code development platforms.
Smart Images

Figure CN121615136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software development security technology, and more specifically, to a software security vulnerability detection method, apparatus, device, and storage medium. Background Technology
[0002] With the deepening of digital transformation, low-code development platforms, with their efficient development mode of graphical drag-and-drop and lightweight configuration, have become core tools for rapidly building applications in fields such as e-commerce, enterprise OA, and government services. However, low-code projects generally suffer from low automation in dependency management and a lack of security testing processes, resulting in a much higher attack rate on dependencies-related vulnerabilities than traditional code projects, making them a key target for cyberattacks.
[0003] Currently, mainstream third-party dependency security detection methods in the industry primarily rely on Static Code Analysis (SAST) tools (such as SonarQube) or Software Composition Analysis (SCA) tools (such as Snyk and Sonatype Nexus). The core logic of these tools is to match the name and version information of the dependency to be detected with static vulnerability databases such as CVEs (Common Vulnerability Disclosures) and NVDs (National Vulnerability Database) to identify known vulnerabilities. However, this version-based matching method lacks risk analysis considering the actual scenario of low-code projects (such as dependency usage, deployment environment, and business importance), ultimately leading to a high false positive rate and failing to meet the needs of precise security control. Summary of the Invention
[0004] The problem addressed by this invention is how to reduce the false alarm rate of security detection.
[0005] To address the above problems, this invention provides a software security vulnerability detection method, apparatus, device, and storage medium.
[0006] In a first aspect, the present invention provides a software security vulnerability detection method, comprising: When a security check is automatically triggered based on preset conditions, the dependency declaration file of the software project to be checked is obtained, and the dependency declaration file is parsed to obtain a list of dependencies to be checked. Obtain structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base; Based on a pre-defined large language model, a risk detection and assessment report is obtained according to the structured security knowledge fragments, the list of dependencies to be checked, and the scenario data of the software project to be tested.
[0007] Optionally, parsing the dependency declaration file to obtain a list of dependencies to be checked includes: Identify the type of the dependency declaration file; The dependency declaration file is parsed using the corresponding parsing algorithm to obtain the corresponding dependency information. According to the preset JSON format specification, the dependency information is formatted and converted to generate a list of dependencies to be checked with uniform fields.
[0008] Optionally, the list of dependencies to be checked includes multiple dependency libraries and their corresponding version information; the step of obtaining structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base includes: Generate a separate query request for each dependent library and its corresponding version information. Based on the name of the dependent library in each independent query request, a matching and filtering process is performed in the dynamic security knowledge base to obtain preliminary matching records; Based on the ecosystem to which the independent query request belongs, the preliminary matching records are filtered to obtain name-ecosystem matching records; the name-ecosystem matching records include affected version range information; Based on semantic versioning rules, the version information in each independent query request is compared with the corresponding affected version range information to obtain the structured security knowledge fragment.
[0009] Optionally, the construction process of the dynamic security knowledge base includes: Unstructured and semi-structured data are obtained through a pre-set multi-source database; All the unstructured data and semi-structured data are transformed to obtain initial structured knowledge, and the dynamic security knowledge base is constructed based on the initial structured knowledge. The dynamic security knowledge base adopts a hybrid storage architecture of relational database and vector database. The dynamic security knowledge base is dynamically updated based on preset triggering conditions.
[0010] Optionally, the preset triggering conditions include timed synchronization update mechanism triggering conditions and real-time update triggering conditions, wherein the real-time update triggering conditions include high-risk vulnerability event triggering and emergency security announcement triggering.
[0011] Optionally, the risk detection and assessment report includes vulnerability impact assessment results, comprehensive risk level assessment results, remediation recommendation results, and license compliance check results; after obtaining the risk detection and assessment report based on the structured security knowledge fragment and the list of dependencies to be checked, it further includes: The vulnerability impact assessment results, the comprehensive risk level assessment results, the remediation suggestion results, and the license compliance check results are structured and transformed to obtain JSON format data; The risk detection and assessment report is output through a visual interface, and the data is output in JSON format via a structured data interface.
[0012] Optionally, after obtaining the risk detection and assessment report based on the structured security knowledge fragment and the list of dependencies to be inspected, the process further includes: According to the automatic quality access control rules set for the software project to be tested, the risk detection and assessment report and the corresponding JSON format data are integrated into the low-code development platform of the software project to be tested. The risk level in the risk detection and assessment report is determined by the automatic quality access control rules, and the corresponding execution strategy is obtained.
[0013] Secondly, the present invention provides a software security vulnerability detection device, comprising: The acquisition unit is used to acquire the dependency declaration file of the software project to be tested when a security check is automatically triggered based on preset conditions, and to parse the dependency declaration file to obtain a list of dependencies to be checked. The filtering unit is used to obtain structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base. The detection unit is used to generate a risk detection and assessment report based on a preset large language model, the structured security knowledge fragments, the list of dependencies to be checked, and the scenario data of the software project to be tested.
[0014] Thirdly, the present invention provides a software security vulnerability detection device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the software security vulnerability detection method as described in the first aspect.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the software security vulnerability detection method as described in the first aspect.
[0016] The beneficial effects of the software security vulnerability detection method, apparatus, system, and storage medium of the present invention are: By automatically triggering security checks under preset conditions, the system automates vulnerability detection, reducing manual intervention and improving detection timeliness and efficiency. Parsing dependency declaration files to obtain a list of dependencies to be checked ensures that the detection scope accurately focuses on dependencies actually used in the project, avoiding invalid detection. Relying on a dynamic security knowledge base to obtain relevant structured security knowledge fragments guarantees the timeliness and accuracy of vulnerability information, providing a reliable basis for risk assessment. Combining a large language model with structured security knowledge, dependency lists, and project scenario data to generate assessment reports enables scenario-based and precise risk analysis, effectively reducing false positive rates and improving the relevance and practicality of reports. Ultimately, this comprehensively enhances the automation, precision, and intelligence of software project dependency security management. Attached Figure Description
[0017] Figure 1 This is one of the flowcharts of a software security vulnerability detection method according to an embodiment of the present invention; Figure 2 This is a schematic diagram showing the security detection results of an embodiment of the present invention in the user interface of a low-code platform; Figure 3 This is a second flowchart illustrating a software security vulnerability detection method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a software security vulnerability detection device according to an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0020] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0021] It should be noted that the terms "one" and "more" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0023] The shortcomings of existing detection methods are mainly reflected in four aspects: First, they have poor scenario adaptability and can only perform vulnerability matching based on dependency versions. They cannot analyze risks in combination with the actual scenario of low-code projects (such as whether the dependency is used for public network payment modules or local testing tools, whether the project is a core transaction system or an internal demonstration application), resulting in a high false positive rate. For example, they may judge dependency vulnerabilities used only for local development as high-risk vulnerabilities in the production environment, or ignore the potential impact of medium-risk vulnerabilities in core business modules. Second, the timeliness and completeness of knowledge are insufficient. Vulnerability databases mostly adopt weekly full updates, which cannot synchronize with the emergency vulnerabilities in open source project security announcements (such as GitHub Security Advisories) in real time (such as the log4j "Log4j2" vulnerability, which was disclosed and then attacked on a large scale in only 2 hours). Furthermore, license compliance data is not integrated, which may easily overlook compliance risks such as commercial closed-source projects using GPL license dependencies. Third, the assessment and recommendations are too generalized, only able to output vague conclusions such as the existence of CVE-XXXX vulnerabilities and recommendations to upgrade to the latest version, without providing accurate security versions or actionable remediation steps, which contradicts the demand for low-code and efficient development; Fourth, the platform has low integration, and most of them are independently deployed tools. Developers need to manually upload dependency lists and download reports. They cannot be seamlessly integrated with the development environment of low-code platforms and CI / CD pipelines, which makes security detection a post-event remedy and significantly increases the cost of vulnerability repair.
[0024] In summary, existing third-party security detection methods suffer from problems such as lack of scenario awareness, lagging knowledge updates, generalized suggestions, and poor integration, making it difficult to adapt to the needs of low-code development platforms for efficient, accurate, and automated security management. There is an urgent need for a software security vulnerability detection solution that combines project scenarios, real-time knowledge, and low-code processes to balance development efficiency and security reliability.
[0025] To address the problems existing in the aforementioned related technologies, embodiments of the present invention provide a software security vulnerability detection method, apparatus, device, and storage medium.
[0026] like Figure 1 As shown in the figure, an embodiment of the present invention provides a software security vulnerability detection method, including: Step S100: When a security check is automatically triggered based on preset conditions, the dependency declaration file of the software project to be checked is obtained, and the dependency declaration file is parsed to obtain a list of dependencies to be checked.
[0027] Specifically, this step is the initiation and basic data preparation stage of software security vulnerability detection, and its core includes two sub-processes: security check triggering and dependency declaration file parsing. Automatic triggering of security checks: The security check process is automatically initiated when preset conditions are met (such as a low-code platform user submitting a project, the CI / CD pipeline executing a build command, or the scheduled detection time being reached), without the need for manual triggering. For example, the check will be triggered when a developer clicks the deployment preview button on the low-code platform, or when the scheduled task starts at 3:00 AM every day.
[0028] Among them, CI / CD pipeline (Continuous Integration / Continuous Deployment Pipeline) is a core practice in software development used to automate code integration, testing, and deployment, achieving rapid iteration and reliable delivery through containerization technology and automation toolchain.
[0029] Dependency declaration file acquisition and parsing: Automatically scan the root directory and subdirectories of the software project to be tested, identify and obtain all dependency declaration files (such as package.json for front-end projects, pom.xml for Java projects, build.gradle for Android projects, etc.); parse each dependency declaration file to extract a list of dependencies to be checked, which may include key information such as dependency library name, version number, dependency type (production / development environment), providing a data foundation for subsequent testing processes.
[0030] By automatically triggering security checks based on preset conditions, the system effectively avoids the problems of manual omissions or detection lagging behind the development schedule. This ensures that security checks are synchronized with the project development process (such as immediate detection after code submission), bringing the vulnerability discovery point forward from the production deployment stage to the development stage. This significantly improves the timeliness and automation of detection and reduces the cost of vulnerability remediation. At the same time, by parsing the actual dependency declaration files of the project, the system only detects the actual dependencies that are actually introduced, eliminating unused redundant dependencies or false declarations. This avoids the resource waste and inefficiency caused by traditional tools scanning all irrelevant dependencies, ensuring the accuracy of the detection scope.
[0031] Step S200: Obtain structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base.
[0032] Specifically, step S200 is the core step connecting the dependencies to be checked with security knowledge. It aims to accurately extract structured security information related to the list of dependencies to be checked from a pre-set dynamic security knowledge base. The specific process is as follows: Using the list of dependencies to be checked generated in step S100 as the retrieval basis, a query is automatically initiated to the dynamic security knowledge base for each dependency in the list (such as a third-party library with a specific name and version). The dynamic security knowledge base integrates multi-dimensional security data such as real-time updated vulnerability information, security announcements, and compliance clauses. By matching the dependency's identification information (such as name and version) in the knowledge base, security content directly related to the dependency (such as known vulnerability details, risk level, license type, etc.) is filtered out. Subsequently, this scattered security information can be transformed into structured knowledge fragments with a unified format and clearly defined fields (such as structured data containing fixed fields such as "vulnerability number, risk level, scope of impact, and compliance status"), ultimately forming a structured security knowledge set that corresponds one-to-one with the list of dependencies to be checked.
[0033] By matching the dynamic security knowledge base with the list of dependencies to be checked, firstly, it ensures that the acquired security knowledge is directly related to the actual dependencies of the project, avoiding interference from irrelevant information and improving the targeting of subsequent analysis; secondly, the real-time updating feature of the dynamic security knowledge base ensures the timeliness of security knowledge, enabling the timely inclusion of newly disclosed vulnerabilities or compliance requirements, solving the problem of missed detection caused by the lag in information from traditional static databases; thirdly, the output of structured knowledge fragments provides standardized input for downstream processes (such as risk assessment using large language models), reducing the cost of data format conversion and improving the efficiency of the overall process; finally, by integrating multi-dimensional knowledge such as vulnerabilities and compliance, comprehensive coverage of security risks of dependencies is achieved, avoiding the limitations of single-dimensional detection and laying a complete and reliable knowledge foundation for the subsequent generation of a comprehensive risk assessment report.
[0034] Step S300: Based on the preset large language model, a risk detection and assessment report is obtained according to the structured security knowledge fragments, the list of dependencies to be checked, and the scenario data of the software project to be tested.
[0035] Specifically, the list of dependencies to be checked obtained in step S100 (including dependency name, version, function, etc.), the structured security knowledge fragments extracted in step S200 (including vulnerability details, risk level, compliance, etc.), and the scenario data of the software project to be tested (such as the project type being an e-commerce transaction system, the deployment environment being a production public network, and the core business module being a payment process, etc.) are used as joint inputs and fed into a large language model optimized in the field of software security. The model can deeply understand the relationship between the three, for example, by combining the security knowledge that log4j-core@2.14.1 has a high-risk RCE vulnerability, the list information of the dependency used for logging in the payment module, and the scenario data of the project processing sensitive payment data and being deployed on the public network, to perform comprehensive reasoning and finally generate a risk detection and assessment report that can include vulnerability impact assessment (such as the specific threat to the payment business), risk level assessment (such as the actual risk level after adjustment based on the scenario), remediation suggestions (such as specific upgrade steps adapted to the project environment), and license compliance judgment, etc.
[0036] By integrating multi-source data through a large language model for analysis, the risk assessment is firstly made more scenario-based and precise, avoiding the limitations of traditional tools that only rate vulnerabilities based on the vulnerability itself. For example, the same medium-risk vulnerability may be assigned different actual risk weights in an internal OA system and an e-commerce payment system, making the assessment results more aligned with the actual security needs of the project. Secondly, the natural language understanding and generation capabilities of the large language model can transform structured data into easy-to-understand and logically coherent assessment reports, including descriptions of specific business impacts and actionable remediation steps (e.g., for the log4j vulnerability in the payment module, it is recommended to upgrade to version 2.17.0 and execute command XXX), reducing the understanding and execution costs for developers. Thirdly, by integrating dependency lists, security knowledge, and scenario data, a full-link analysis of vulnerability-dependency-business is achieved, ensuring that the report not only points out the problem but also explains the cause of the risk and the possible business consequences, providing project teams with complete decision support from risk discovery to risk resolution. Finally, the flexibility of the large language model allows it to adapt to the different scenarios of different types of projects (such as government systems and financial platforms), improving the universality and practicality of the detection method.
[0037] In this embodiment, security checks are automatically triggered by preset conditions, achieving automated initiation of vulnerability detection, reducing manual intervention, and improving detection timeliness and efficiency. By parsing dependency declaration files to obtain a list of dependencies to be checked, the detection scope is ensured to accurately focus on dependencies actually used in the project, avoiding invalid detection. Relying on a dynamic security knowledge base to obtain relevant structured security knowledge fragments ensures the timeliness and accuracy of vulnerability information, providing a reliable basis for risk assessment. Combining a large language model with structured security knowledge, dependency lists, and project scenario data to generate an assessment report enables scenario-based and precise risk analysis, effectively reducing false positive rates, improving the report's relevance and practicality, and ultimately comprehensively enhancing the automation, precision, and intelligence of software project dependency security management.
[0038] Optionally, parsing the dependency declaration file to obtain a list of dependencies to be checked includes: Identify the type of the dependency declaration file; The dependency declaration file is parsed using the corresponding parsing algorithm to obtain the corresponding dependency information. According to the preset JSON format specification, the dependency information is formatted and converted to generate a list of dependencies to be checked with uniform fields.
[0039] Specifically, the parsing and standardization of dependency declaration files aims to transform dependency information in different formats into a unified and usable checklist, which includes three steps: File type identification: Scan the obtained dependency declaration files and automatically identify their type (such as package.json for front-end projects, pom.xml for Java projects, build.gradle for Android projects, etc.) by file extension (such as .json, .xml, .gradle), header identifiers or syntax features, and determine the corresponding development language or build tool of the file.
[0040] Targeted parsing and processing: Based on the identified file type, the appropriate parsing algorithm is invoked (such as a JSON parser for package.json, an XML parser for pom.xml, and a specific syntax parser for build.gradle scripts) to extract the core information of the dependencies from the file, including but not limited to the dependency library name, version number, version constraint rules (such as "^1.2.0", ">=2.0.0"), dependency type (production environment / development environment), source repository, etc.
[0041] Standardized format conversion: The extracted dependency information is uniformly converted according to the preset JSON format specification. For example, information such as name-version-type is mapped to fixed fields (such as "name": "log4j-core", "version": "2.14.1", "type": "production"), and finally a list of dependencies to be checked with uniform fields and consistent structure is generated.
[0042] This process ensures accurate parsing of dependency declaration files in different formats by first identifying the file type and then calling the corresponding parsing algorithm, avoiding omissions or errors in information extraction due to format differences. At the same time, the standardized processing based on the preset JSON format specification transforms the scattered dependency information into a unified structure, providing consistent input for subsequent matching with dynamic security knowledge bases and risk assessment of large language models. This reduces the data adaptation costs of downstream processes, improves the efficiency and compatibility of the overall detection process, and accurately preserves the key attributes of dependencies (such as version constraints and environment types), providing complete basic data support for subsequent risk analysis.
[0043] Optionally, the list of dependencies to be checked includes multiple dependency libraries and their corresponding version information; the step of obtaining structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base includes: Generate a separate query request for each dependent library and its corresponding version information. Based on the name of the dependent library in each independent query request, a matching and filtering process is performed in the dynamic security knowledge base to obtain preliminary matching records; Based on the ecosystem to which the independent query request belongs, the preliminary matching records are filtered to obtain name-ecosystem matching records; the name-ecosystem matching records include affected version range information; Based on semantic versioning rules, the version information in each independent query request is compared with the corresponding affected version range information to obtain the structured security knowledge fragment.
[0044] Specifically, the core process of accurately extracting security knowledge related to dependencies from a dynamic security knowledge base involves four steps to complete knowledge matching for each dependency library and version information in the list of dependencies to be checked: Generate independent query requests: For each dependency library in the list (such as "log4j-core" and "lodash") and its corresponding version information (such as "2.14.1" and "4.17.19"), automatically generate an independent query request containing the dependency library name, version information and its ecosystem (such as Java and Node.js), ensuring that each dependency can be retrieved individually.
[0045] Name matching filtering: Based on the dependency library names in the query request, a preliminary match is performed in the dynamic security knowledge base to filter out all security records with the same name (such as all vulnerability information containing "log4j-core"), forming a preliminary matching record.
[0046] Ecosystem filtering: Combine the ecosystem (e.g., Java) in the query request to perform secondary filtering on the initial matching records, exclude dependency records with the same name across ecosystems (e.g., remove irrelevant records named "log4j-core" in the Python ecosystem), and obtain records that match both the name and the ecosystem. These records contain information on the affected version range of the dependency library (e.g., "log4j-core 2.0-beta9 to 2.14.1 have vulnerabilities").
[0047] Version range comparison: Based on semantic versioning rules (such as SemVer), the specific version information in the query request (such as "2.14.1") is precisely compared with the affected version range in the matching records to determine whether the version is within the scope of the vulnerability. Finally, the records that meet the conditions (such as vulnerability number, risk level, remediation suggestions, etc.) are integrated into structured security knowledge fragments.
[0048] By employing a multi-layered verification mechanism combining independent querying, name matching, ecosystem filtering, and version comparison, the system first resolves the confusion issue of dependencies with the same name across different ecosystems (e.g., distinguishing between "log4j" for Java and Python), ensuring the ecosystem accuracy of the matching records. Second, precise comparison based on semantic versioning rules avoids misjudgments caused by differences in version descriptions (e.g., "^2.0" vs. "2.1.x"), significantly improving the accuracy of vulnerability matching. Simultaneously, generating independent queries for each dependency and progressively filtering them ensures that structured security knowledge fragments contain only information directly related to the dependency being checked, reducing interference from irrelevant data and providing high-quality knowledge input for subsequent risk assessment. Ultimately, this enhances the reliability and efficiency of the overall detection process.
[0049] Optionally, the construction process of the dynamic security knowledge base includes: Unstructured and semi-structured data are obtained through a pre-set multi-source database; All the unstructured data and semi-structured data are transformed to obtain initial structured knowledge, and the dynamic security knowledge base is constructed based on the initial structured knowledge. The dynamic security knowledge base adopts a hybrid storage architecture of relational database and vector database. The dynamic security knowledge base is dynamically updated based on preset triggering conditions.
[0050] Optionally, the preset triggering conditions include timed synchronization update mechanism triggering conditions and real-time update triggering conditions, wherein the real-time update triggering conditions include high-risk vulnerability event triggering and emergency security announcement triggering.
[0051] Specifically, the mechanism for building and updating the dynamic security knowledge base aims to create a real-time, comprehensive, and structured security knowledge reserve, which includes three core components: Multi-source data acquisition: Data is collected from pre-defined multi-source databases (such as the CVE vulnerability database, NVD database, open-source project security bulletin platform, license compliance database, etc., through API interfaces, web crawling tools, or data synchronization protocols. Among them, CVE (Common Vulnerabilities and Exposures): provides a unified standard for security vulnerability identification, assigning a unique number (such as CVE-2021-44428) to each known vulnerability, including vulnerability description, scope of impact, and reference information. NVD (National Vulnerability Database): is a vulnerability database that expands upon the CVE list, including risk scores, remediation suggestions, and other detailed information) through the CVE database. This includes unstructured data (such as vulnerability bulletin text described in natural language and security analysis reports) and semi-structured data (such as tagged XML format vulnerability information and JSON format license terms).
[0052] Data structuring and storage: The collected unstructured and semi-structured data are transformed and processed. For example, key information such as CVE number, scope of impact, and risk level in vulnerability announcements are extracted using natural language processing technology. Semi-structured XML data is converted into key-value pair format to form initial structured knowledge containing unified fields. Subsequently, a hybrid storage architecture of relational database + vector database is adopted. The relational database stores structured attributes (such as vulnerability number and version range), and the vector database stores semantic vectors (such as semantic features of vulnerability descriptions), taking into account both precise query and semantic retrieval needs.
[0053] Dynamic update triggering: The knowledge base is automatically updated based on preset conditions, including a timed synchronization mechanism (such as synchronizing the CVE database every 6 hours and synchronizing open source project announcements daily) and real-time update triggering conditions (such as immediately synchronizing when a high-risk vulnerability with a CVSS score ≥ 9.0 is detected and released, and triggering an update when an open source project releases a security announcement marked with an emergency fix), ensuring the timeliness of knowledge.
[0054] By collecting data from multiple sources, the system covers security information across multiple dimensions, including vulnerabilities and compliance, thus avoiding the limitations of a single data source. After structured processing, the knowledge is stored in a unified format. Combined with a hybrid storage architecture, it supports both efficient and precise matching (such as querying vulnerabilities in a specific version) and semantic association retrieval (such as finding remediation solutions for similar vulnerabilities). The dynamic update mechanism of "timed + real-time" ensures both the periodic updating of regular knowledge and the second-level response to critical information such as high-risk vulnerabilities and emergency announcements (such as updating the knowledge base within 10 minutes after the disclosure of a log4j vulnerability). This ensures that downstream detection processes are always based on the latest knowledge, significantly reducing the risk of missed detections due to information lag, and providing a comprehensive, timely, and reliable knowledge foundation for overall security detection.
[0055] Optionally, the risk detection and assessment report includes vulnerability impact assessment results, comprehensive risk level assessment results, remediation recommendation results, and license compliance check results; after obtaining the risk detection and assessment report based on the structured security knowledge fragment and the list of dependencies to be checked, it further includes: The vulnerability impact assessment results, the comprehensive risk level assessment results, the remediation suggestion results, and the license compliance check results are structured and transformed to obtain JSON format data; The risk detection and assessment report is output through a visual interface, and the data is output in JSON format via a structured data interface.
[0056] Specifically, for the four core results in the risk detection and assessment report (vulnerability impact assessment results, such as the log4j vulnerability in the e-commerce payment module potentially leading to sensitive data leakage; comprehensive risk level assessment results, such as high risk; remediation recommendation results, such as upgrading to log4j-core@2.17.0; and license compliance check results, such as the Apache-2.0 license being compatible with the project's MIT license), a structured mapping is performed according to preset field specifications (such as "vulnerability_impact", "risk_level", "remediation", "license_compliance"). The results described in natural language are transformed into clear key-value pair JSON format data, ensuring that each result dimension has corresponding standardized fields. For example, {"vulnerability_impact": "The log4j vulnerability in the payment module can lead to the leakage of user bank card numbers","risk_level": "high risk","remediation": "upgrade to log4j-core@2.17.0, execute mvn clean install-U","license_compliance": "compatible"}.
[0057] Among them, vulnerability_impact: the vulnerability impact assessment result, is used to describe the specific harm or business impact that the vulnerability may cause to the software project, such as "the log4j vulnerability in the payment module may lead to remote attackers stealing users' bank card information" and "the access control dependency vulnerability may lead to the risk of unauthorized access".
[0058] risk_level: The overall risk level assessment result, used to identify the overall risk level of the vulnerability (usually divided into high risk, medium risk, low risk, etc.). It is the result of a comprehensive judgment based on factors such as the harm of the vulnerability itself and the project scenario (such as whether it involves core business). For example, "high risk" and "medium risk".
[0059] Remediation: The remediation suggestion results are used to provide specific solutions for the vulnerability. They usually include actionable solutions such as version upgrade suggestions, alternative dependency recommendations, and configuration modification steps. For example, to upgrade to log4j-core@2.17.0, execute the command: mvn clean package -Dlog4j.version=2.17.0.
[0060] license_compliance: License compliance check result, used to indicate whether the license type of the dependency is compatible with the main license of the project, to avoid legal risks caused by license conflicts, such as "compatible (dependency is Apache-2.0 license, main project license is MIT)" or "incompatible (dependency is GPLv3 license, project is commercial closed source).
[0061] Multiple output formats: On the one hand, the risk detection and assessment report is presented in a visual interface using charts, lists, and text (e.g., using red alerts to identify high-risk vulnerabilities, and using tables to display dependency names, risk levels, and remediation recommendations). Figure 2 As shown, the security detection results (data included in the risk detection assessment report) are displayed in the user interface of the low-code platform, making it easy for developers to view them intuitively. On the other hand, JSON format data is output through a preset structured data interface (such as RESTAPI) for use by automated platforms such as CI / CD tools (such as Jenkins) and project management systems (such as Jira) to support subsequent process control (such as CI / CD pipelines using the "risk_level" field in JSON to determine whether to block the build).
[0062] By converting the assessment results into JSON format data, machine-readable results are achieved, providing a standardized interface for integration with automated systems such as CI / CD and project management. This avoids the inefficiency and errors of manual report parsing (e.g., CI / CD tools can directly read the "risk_level" field in the JSON and automatically execute blocking logic). Simultaneously, the visual interface output allows non-technical personnel to clearly understand the risk situation (e.g., project managers can quickly grasp the number and core impact of high-risk vulnerabilities through a dashboard), balancing the needs of both "automated integration" and "human decision-making." Furthermore, the structured conversion ensures the uniformity of fields and the standardization of format in the assessment results. Whether for cross-system transmission or long-term archiving, data consistency is maintained, reducing information loss or misunderstanding caused by formatting issues, ultimately improving the practicality and adaptability of the risk assessment results.
[0063] Optionally, after obtaining the risk detection and assessment report based on the structured security knowledge fragment and the list of dependencies to be inspected, the process further includes: According to the automatic quality access control rules set for the software project to be tested, the risk detection and assessment report and the corresponding JSON format data are integrated into the low-code development platform of the software project to be tested. The risk level in the risk detection and assessment report is determined by the automatic quality access control rules, and the corresponding execution strategy is obtained.
[0064] Specifically, the closed-loop management process, which deeply integrates risk assessment results with low-code development processes, aims to automate the handling of security risks. This process involves two steps: Automatic quality access control rule settings and data integration: Based on the attributes of the software project to be tested (such as the project type being a government core system, e-commerce transaction platform, etc.), automatic quality access control rules are pre-configured. For example, for government systems, "deployment will be blocked when high-risk vulnerabilities exist" and "manual review will be triggered when the number of medium-risk vulnerabilities exceeds 3" can be set. For e-commerce platforms, "payment module vulnerabilities, regardless of level, must be fixed immediately" can be set. At the same time, the risk detection and assessment report (including vulnerability details, remediation suggestions, etc.) and the corresponding JSON format data (such as fields such as "risk_level" and "vulnerability_impact") are integrated into the low-code development platform through an interface, so that the platform can directly read the risk data and associate it with the project development process (such as displaying risk warnings on the "deployment application" page).
[0065] Risk level assessment and execution strategy generation: The low-code platform calls the automatic quality access control rules to match and determine the risk level in the risk detection and assessment report. For example, if the payment module of an e-commerce project has a "high-risk" vulnerability, and the rule is set to "high-risk vulnerability in core module → block deployment and notify the development team", the system automatically generates the execution strategy: immediately suspend the current deployment process, push vulnerability details and remediation suggestions to the developers through the platform message center, and mark the project board with a "security blocked" status; if the risk level is "low-risk" and the rule allows "continue deployment after recording the risk", the strategy is "allow deployment and record the risk to the project security profile simultaneously".
[0066] By customizing automated quality access control rules for different projects, scenario-based adaptation of security management is achieved, avoiding a "one-size-fits-all" approach to risk handling (such as adopting differentiated management strategies for core systems and internal tools). Integrating assessment results with the low-code platform allows risk information to be directly embedded in the development process (such as real-time alerts during deployment), solving the problem of "reports being disconnected from the development environment" in traditional tools. Based on rules, risk levels are automatically determined and execution strategies are generated, ensuring that high-risk risks are blocked in a timely manner (such as preventing the payment module vulnerability from going live) while reducing the interference of low-risk risks on development efficiency (such as allowing recording before continuing). Ultimately, an automated closed loop of "detection-assessment-handling" is formed, ensuring software security while also addressing the core requirement of "efficient iteration" in low-code development.
[0067] In some embodiments, such as Figure 3 As shown, the software security vulnerability detection method is applied to a security inspection system integrated with a RAG architecture (where RAG (Retrieval-Augmented Generation) is an artificial intelligence technology that combines retrieval and generation to enhance the knowledge coverage and response accuracy of the model). The integrated RAG architecture security inspection system includes a knowledge base construction module: collecting and structurally storing security knowledge such as CVEs, NVDs, open-source license information, and code behavior characteristics; a query processing module: parsing dependency declaration files (such as package.json and pom.xml) in low-code platforms and extracting dependency information; a RAG retrieval and generation module: retrieving information from the knowledge base based on dependency information and generating context-sensitive risk assessment reports and remediation suggestions using a large language model; and a result output and integration module: presenting the inspection results in a visual manner and supporting integration with CI / CD processes. The specific steps of the software security vulnerability detection method are as follows: Step 1: Trigger the check: When the low-code platform (including the user interface and CI / CD Pipeline) meets preset conditions (such as when a user submits a project or when the CI / CD pipeline executes a build command), it triggers a security check process and sends the check request to the query processing module.
[0068] Step 2: Dependency query and retrieval request generation: After receiving the trigger request, the query processing module obtains the dependency declaration file of the software project to be tested, parses it to obtain the list of dependencies to be checked, and passes the dependency list to the RAG retrieval and generation module to generate a retrieval request for the dependency list.
[0069] Step 3: Safety Knowledge Retrieval: The RAG retrieval and generation module sends retrieval requests to the security knowledge base, retrieving structured security knowledge fragments related to the list of dependencies to be checked. The security knowledge base is dynamically updated through a data synchronization mechanism with external data sources (CVE, NVD, license databases).
[0070] Step 4: Generate a risk detection and assessment report: The RAG retrieval and generation module combines structured security knowledge fragments obtained from the security knowledge base, a list of dependencies to be checked, and scenario data of the software projects to be tested to generate a risk detection and assessment report, which is then passed to the results output and integration module.
[0071] Step 5: Result Output and Process Integration: The results output and integration module visualizes the risk detection and assessment report (such as displaying it on the low-code platform user interface) and integrates structured data into the low-code platform's CI / CD Pipeline, supporting automated process control based on report results (such as quality access control decisions, process blocking, or continuation).
[0072] Step 6: Return Results and Close the Process Loop: The final results are returned to the low-code platform, completing the entire closed loop from "triggering inspection" to "result application", and realizing deep linkage between software security vulnerability detection and low-code development process.
[0073] like Figure 4 As shown in the figure, an embodiment of the present invention provides a software security vulnerability detection device, comprising: The acquisition unit is used to acquire the dependency declaration file of the software project to be tested when a security check is automatically triggered based on preset conditions, and to parse the dependency declaration file to obtain a list of dependencies to be checked. The filtering unit is used to obtain structured security knowledge fragments related to the list of dependencies to be checked based on a preset dynamic security knowledge base. The detection unit is used to generate a risk detection and assessment report based on a preset large language model, the structured security knowledge fragments, the list of dependencies to be checked, and the scenario data of the software project to be tested.
[0074] This invention provides a software security vulnerability detection device, including a memory and a processor; the memory is used to store a computer program; the processor is used to implement the software security vulnerability detection method described above when the computer program is executed.
[0075] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the software security vulnerability detection method described above.
[0076] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A software security vulnerability detection method, characterized by, The application relates to a method for automatically triggering a security check based on a preset condition, and comprises the following steps: When the security check is automatically triggered based on the preset condition, a dependent declaration file of a software item to be detected is acquired, and the dependent declaration file is parsed to obtain a list of dependent items to be checked; According to a preset dynamic security knowledge base, a structured security knowledge fragment related to the list of dependent items to be checked is obtained; Based on a preset large language model, a risk detection evaluation report is obtained according to the structured security knowledge fragment, the list of dependent items to be checked and scene data of the software item to be detected.
2. The software security vulnerability detection method of claim 1, wherein, The parsing of the dependent declaration file to obtain the list of dependent items to be checked comprises the following steps: The type of the dependent declaration file is identified; According to a corresponding parsing algorithm, the dependent declaration file is parsed to obtain corresponding dependent item information, According to a preset JSON format specification, the dependent item information is format-converted to generate the list of dependent items to be checked with uniform fields.
3. The software security vulnerability detection method of claim 2, wherein, The list of dependent items to be checked comprises a plurality of dependent libraries and corresponding version information; the structured security knowledge fragment related to the list of dependent items to be checked is obtained according to a preset dynamic security knowledge base, and comprises the following steps: According to each dependent library and corresponding version information, a corresponding independent query request is generated, According to the name of the dependent library in each independent query request, the dynamic security knowledge base is matched and screened to obtain preliminary matching records; Based on the belonging ecosystem in the independent query request, the preliminary matching records are filtered to obtain name-ecosystem matching records; the name-ecosystem matching records comprise affected version range information; According to a semantic version rule, the version information in each independent query request is compared with the corresponding affected version range information to obtain the structured security knowledge fragment.
4. The software security vulnerability detection method of claim 1, wherein, The construction process of the dynamic security knowledge base comprises the following steps: Unstructured data and semi-structured data are acquired through a preset multi-source database; All the unstructured data and the semi-structured data are converted to obtain initial structured knowledge, and the dynamic security knowledge base is constructed according to the initial structured knowledge; the dynamic security knowledge base adopts a hybrid storage architecture of a relational database and a vector database; The dynamic security knowledge base is dynamically updated based on a preset trigger condition.
5. The software security vulnerability detection method of claim 4, wherein, The preset trigger condition comprises a timing synchronization update mechanism trigger condition and a real-time update trigger condition; the real-time update trigger condition comprises a high-risk vulnerability event trigger and an emergency security announcement trigger.
6. The method of claim 1, wherein, The risk detection evaluation report comprises a vulnerability impact evaluation result, a risk level comprehensive evaluation result, a repair suggestion result and a license compliance check result; After the risk detection evaluation report is obtained according to the structured security knowledge fragment and the list of dependent items to be checked, the following steps are further included: The vulnerability impact evaluation result, the risk level comprehensive evaluation result, the repair suggestion result and the license compliance check result are structure-converted to obtain JSON format data; The risk detection evaluation report is output through a visual interface, and the JSON format data is output in a structured data interface.
7. The software security vulnerability detection method of claim 6, wherein, The risk detection evaluation report is obtained according to the structured security knowledge segment and the dependent item list to be checked. According to the corresponding automatic quality access control rule set by the software project to be detected, the risk detection evaluation report and the corresponding JSON format data are integrated into the low-code development platform of the software project to be detected. The risk level in the risk detection evaluation report is judged by the automatic quality access control rule, and a corresponding execution strategy is obtained.
8. A software security vulnerability detection apparatus characterized by comprising: It comprises: An acquisition unit is configured to acquire a dependent declaration file of a software project to be detected and parse the dependent declaration file to obtain a dependent item list to be checked when a security check is automatically triggered based on a preset condition; A screening unit is configured to obtain a structured security knowledge segment related to the dependent item list to be checked according to a preset dynamic security knowledge base; A detection unit is configured to obtain a risk detection evaluation report based on a preset large language model according to the structured security knowledge segment, the dependent item list to be checked and scene data of the software project to be detected.
9. A software security vulnerability detection device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program comprises instructions for causing the processor to perform the steps of: receiving a software program; determining a plurality of software security vulnerabilities in the software program; and outputting a report of the plurality of software security vulnerabilities in the software program. The processor executes the computer program to implement the software security vulnerability detection method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium has a computer program stored thereon, and when the computer program is executed by the processor, the software security vulnerability detection method of any one of claims 1-7 is implemented.