Vulnerability detection method and device, equipment, storage medium and program product
By combining multimodal sample generation and knowledge graphs, this method addresses the issues of weak model generalization and poor cross-modal alignment in existing vulnerability detection methods, achieving efficient, diverse, and robust detection of adversarial examples.
Patent Information
- Application Number
- CN202511064080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-11
AI Technical Summary
Existing vulnerability detection methods rely on limited historical vulnerability data, have weak model generalization ability, cannot effectively associate discrete vulnerability features, and have poor cross-modal alignment, making it difficult to predict combined attacks.
A multimodal sample input generator is used to generate adversarial examples. Feature alignment and cross-modal fusion are achieved through collaborative training of the generator and discriminator. Multimodal retrieval is performed by combining knowledge graphs to detect vulnerability information in adversarial examples.
It improves the diversity of adversarial examples and the robustness of the model in adversarial environments, and enhances the attack chain detection coverage and vulnerability information identification efficiency.
Smart Images

Figure CN120930148A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically to a vulnerability detection method, apparatus, device, storage medium, and program product. Background Technology
[0002] Current cybersecurity threats exhibit two main characteristics: multimodal fusion attacks and the evolution towards stealth. Traditional vulnerability detection methods rely on limited historical vulnerability data, and artificially synthesized samples suffer from poor grammatical compliance and insufficient diversity, resulting in weak model generalization capabilities.
[0003] Existing vulnerability detection technologies mainly employ two approaches: traditional rule-based matching and static analysis methods, and machine learning methods based on mathematical statistical inference. However, these two existing technologies have several technical problems. For example, they cannot effectively associate discrete vulnerability features, resulting in incomplete attack chain detection; at the same time, retrieval-enhanced generation only supports a single modality (such as text), while the semantic space of code and text is misaligned, leading to poor cross-modal alignment; and static segmentation strategies also disrupt the integrity of the code structure, ultimately making it difficult to predict combined attacks. Summary of the Invention
[0004] In view of the above problems, this application provides a vulnerability detection method, apparatus, device, storage medium, and program product.
[0005] According to a first aspect of this application, a vulnerability detection method is provided, comprising: acquiring an original sample, wherein the original sample is a multimodal sample; inputting the multimodal sample into a preset generator, and performing adversarial verification after obtaining an adversarial sample; parsing the adversarial sample, and inputting the extracted entities into a preset knowledge graph to update the knowledge graph; performing multimodal retrieval based on the updated knowledge graph to detect vulnerability information in the adversarial sample; wherein, in the preset generator, the multimodal sample undergoes feature alignment and cross-modal fusion processing.
[0006] According to an embodiment of this application, the vulnerability detection method further includes: adjusting the noise data used by the generator when constructing adversarial samples based on vulnerability information, so as to complete the relearning of the generator.
[0007] According to an embodiment of this application, multimodal samples are input into a preset generator, and adversarial verification is completed after obtaining adversarial samples. This includes: inputting multimodal samples into a preset generator, extracting modal features through a dual-path coding structure and performing feature fusion to obtain adversarial samples; inputting adversarial samples into a preset discriminator, obtaining adversarial loss by performing multi-dimensional verification on the adversarial samples, and sending it back to the generator.
[0008] According to an embodiment of this application, multimodal samples are input into a preset generator, and adversarial verification is completed after obtaining adversarial samples. The method further includes: adjusting the noise injection strategy and routing path selection of the generator according to the multidimensional verification structure of the discriminator.
[0009] According to an embodiment of this application, the knowledge graph is pre-defined, and the pre-defined process includes: acquiring heterogeneous security data from multiple sources, including vulnerability reports; extracting entities and relationships from the heterogeneous security data and storing triples, including attack surface, vulnerability type, and affected components; generating causal relationship chains based on entities and relationships through taint tracing; executing vulnerability attacks according to the causal relationship chains and obtaining attack results; and constructing the knowledge graph after verifying the attack results and determining the attack chain.
[0010] According to an embodiment of this application, the pre-setting process of the knowledge graph further includes: achieving deep fusion of multi-source data by defining heterogeneous edge types of heterogeneous security data and dynamic weight calculation; wherein, the heterogeneous edge types include static data streams, dynamic attack chains and text inference relationships.
[0011] According to an embodiment of this application, the pre-setting process of the knowledge graph further includes: after the knowledge graph is constructed, multi-hop reasoning is implemented based on the relational graph convolutional network to predict potential vulnerability exploitation paths.
[0012] According to an embodiment of this application, parsing adversarial examples and inputting the extracted entities into a preset knowledge graph includes: parsing adversarial examples based on a preset parsing tool to construct an abstract syntax tree structure; traversing the abstract syntax tree structure to extract entities; and inputting the entities into the preset knowledge graph.
[0013] According to embodiments of this application, multimodal retrieval is performed based on knowledge graph updates to detect vulnerability information in adversarial examples, including: determining vulnerability information in adversarial examples through adversarial training and counterfactual reasoning of three-level retrieval; wherein, the three-level retrieval includes sparse retrieval, dense retrieval, and reordering retrieval.
[0014] The second aspect of this application provides a vulnerability detection device, comprising: a raw sample acquisition module for acquiring raw samples, wherein the raw samples are multimodal samples; an adversarial sample generation module for inputting the multimodal samples into a preset generator and performing adversarial verification after obtaining the adversarial samples; a knowledge graph update module for parsing the adversarial samples and inputting the extracted entities into a preset knowledge graph to update the knowledge graph; and a vulnerability information detection module for performing multimodal retrieval based on the updated knowledge graph to detect vulnerability information in the adversarial samples; wherein, in the preset generator, the multimodal samples undergo feature alignment and cross-modal fusion processing.
[0015] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0016] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0017] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description
[0018] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1 This diagram illustrates an application scenario of the vulnerability detection method according to an embodiment of this application.
[0020] Figure 2 A flowchart illustrating a vulnerability detection method according to an embodiment of this application is shown schematically.
[0021] Figure 3 This illustration schematically shows a flowchart of obtaining adversarial samples and completing adversarial verification according to an embodiment of this application;
[0022] Figure 4 A flowchart illustrating the construction of a knowledge graph according to an embodiment of this application is shown schematically;
[0023] Figure 5 This schematically illustrates a flowchart of inputting extracted entities into a knowledge graph according to an embodiment of this application;
[0024] Figure 6 This schematically illustrates a structural block diagram of a vulnerability detection apparatus according to an embodiment of this application; and
[0025] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a vulnerability detection method according to an embodiment of this application. Detailed Implementation
[0026] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0030] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0031] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0032] An embodiment of this application provides a vulnerability detection method, comprising: acquiring original samples, wherein the original samples are multimodal samples; inputting the multimodal samples into a preset generator, and performing adversarial verification after obtaining adversarial samples; parsing the adversarial samples, and inputting the extracted entities into a preset knowledge graph to update the knowledge graph; performing multimodal retrieval based on the updated knowledge graph to detect vulnerability information in the adversarial samples; wherein, in the preset generator, feature alignment and cross-modal fusion processing are performed on the multimodal samples.
[0033] Through the embodiments of this application, the method, by co-training the generator and discriminator, enables the generation of more diverse adversarial examples, effectively covering potential attack surfaces. By adjusting and relearning the generator, not only is the quality of adversarial examples improved, but the robustness of the model in adversarial environments is also enhanced. Furthermore, by combining knowledge graphs and cross-modal retrieval, the method improves attack chain detection coverage, enabling more comprehensive analysis of adversarial examples and allowing vulnerability information within them to be identified with greater efficiency and accuracy.
[0034] Figure 1 The illustration shows an application scenario diagram of the vulnerability detection method according to an embodiment of this application.
[0035] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, and a server 105. Network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0038] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0039] It should be noted that the vulnerability detection method provided in this application embodiment can generally be executed by server 105. Correspondingly, the vulnerability detection device provided in this application embodiment can generally be located in server 105. The vulnerability detection method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the vulnerability detection device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0041] The following will be based on Figure 1 The described scene, through Figures 2-5 The vulnerability detection method according to the embodiments of this application will be described in detail.
[0042] Figure 2 A flowchart illustrating a vulnerability detection method according to an embodiment of this application is shown.
[0043] like Figure 2 As shown, according to an embodiment of this application, the vulnerability detection method specifically includes operations S210 to S240.
[0044] In operation S210, the original samples are obtained. The original samples are multimodal samples.
[0045] For example, raw data from multiple data sources can be obtained. This raw data can be data samples of different modalities such as text, code, or images. These raw data samples include data such as external vulnerability reports and threat intelligence.
[0046] In operation S220, multimodal samples are input into a preset generator, and adversarial verification is completed after adversarial samples are obtained.
[0047] In the embodiments of this application, after obtaining the original samples, high-quality generation and semantic alignment of cross-modal data are automatically achieved through adversarial training of a preset generator and discriminator, effectively solving the problems of sample scarcity and compliance. The process is described in detail below.
[0048] Figure 3 The flowchart illustrating the process of obtaining adversarial samples and completing adversarial verification according to an embodiment of this application is shown in the illustration.
[0049] like Figure 3 As shown, according to an embodiment of this application, multimodal samples are input into a preset generator, and adversarial verification is completed after obtaining adversarial samples. This process specifically includes operations S310 to S320.
[0050] In operation S310, multimodal samples are input into a preset generator, and modal features are extracted and fused through a dual-path coding structure to obtain adversarial samples.
[0051] For example, this generator is a multimodal conditional generator with a dual-channel encoding structure. Depending on the data format, such as a text encoder, the encoding structure utilizes a pre-trained language model to extract semantic features of the text, such as vulnerability descriptions and log information. Long-distance dependencies are captured through the self-attention mechanism of a deep learning architecture, generating context-aware text vectors.
[0052] For example, the encoding structure includes a code encoder that uses graph neural networks to model the syntax and logical relationships of the code based on its structured characteristics (AST, Abstract Syntax Tree), converts the code into a graph structure (nodes are code elements, edges are dependencies), and generates graph embeddings through neighborhood aggregation.
[0053] For example, the encoding structure includes an image encoder that segments an image into a sequence of blocks, extracts global visual features, such as vulnerability diagrams or system architecture diagrams, through self-attention, and inputs them into a deep learning architecture after image segmentation to generate visual vectors aligned with text and code.
[0054] Furthermore, the generator also includes a noise injection layer, which can be used to add controllable noise to the latent space of different levels of the generator to enhance the diversity of generation. The noise injection layer can independently adjust modal features, expand the coverage of generated samples, and decouple different modal features such as code logic and variable naming.
[0055] Furthermore, the generator also includes a dynamic routing mechanism that dynamically selects the generation path based on input features, such as suppressing redundant visual branches in text-dominant tasks to improve computational efficiency.
[0056] It should be noted that the preset generator also performs feature alignment and cross-modal fusion processing on multimodal samples. For example, feature alignment and attention mechanisms can be used to coordinate text, code, and image modal information to form a redundant verification and fault tolerance mechanism.
[0057] Based on the above, multimodal samples are input into a preset generator, and modal features are extracted and fused through a dual-path coding structure to obtain adversarial samples.
[0058] In operation S320, the adversarial sample is input into the preset discriminator. The adversarial loss is obtained by performing multi-dimensional verification on the adversarial sample and then sent back to the generator.
[0059] For example, in cross-modal generation tasks, the discriminator can be a multi-expert discriminator. Through a multi-branch collaborative verification mechanism, it ensures the rationality and cross-modal consistency of the generated samples from three dimensions: syntactic, logical, and semantic alignment. Independent branches judge whether the generated samples (such as code or images) conform to the real data distribution, and verify the matching degree between the generated content and the input conditions through cross-attention, such as whether the code meets the functional requirements of the text description.
[0060] Specifically, the multi-expert discriminator includes three parallel verifiers, such as a code syntax verifier (AST parser), which performs syntax compliance checks. This verification involves parsing the generated code structure using an Abstract Syntax Tree (AST) based on compiler principles. It parses the syntax hierarchy (such as function definitions and loop structures) in a tree structure to verify whether it conforms to the syntax rules of the programming language, detects potential syntax errors (such as unclosed parentheses and illegal operators), and avoids generating code that cannot be compiled or executed.
[0061] For example, this multi-expert discriminator also includes a text logic consistency discriminator to perform logical self-consistency verification. This verification includes checking whether the causal chain and time sequence of the generated text (such as vulnerability descriptions and penetration reports) are reasonable (e.g., whether the attack steps conform to the actual execution process). An inference chain is constructed based on a pre-trained language model, and a loss function is used to quantify the semantic matching degree between the generated text and the input conditions. By comparing the input conditions (such as user requirement text) with the generated content, consistency in functional description and constraints is ensured.
[0062] For example, this multi-expert discriminator also includes a cross-modal alignment detector to achieve multimodal semantic alignment: verifying the semantic consistency of different modal data (such as text descriptions and generated code, images) (e.g., whether the code function matches the text requirements). A cross-modal attention layer is introduced into the discriminator to dynamically associate text keywords with code function blocks / image regions. Adversarial alignment training is completed using an adversarial game between the modality classifier and the feature mapper, forcing the cross-modal representation of the generated samples to have modality invariance.
[0063] Based on the above, it can be understood that the collaborative training process between the generator and discriminator is as follows: after the generator receives multimodal (text, code, image) input, it generates adversarial examples through feature fusion. Then, the discriminator performs multi-dimensional verification (syntax, logic, cross-modal alignment) on the generated adversarial examples, calculates the adversarial loss, and backpropagates it to the generator. Furthermore, after the multi-dimensional verification structure is backpropagated to the generator, the generator's noise injection strategy and routing path selection are adjusted. By adjusting and relearning the generator, the generated adversarial examples can be made more diverse, effectively covering potential attack surfaces.
[0064] Through the embodiments of this application, by co-training the generator and the discriminator, the generated adversarial examples can be made more diverse, effectively covering the potential attack surface and effectively solving the problem of insufficient training data.
[0065] In operation S230, the adversarial examples are parsed, and the extracted entities are input into the preset knowledge graph to complete the knowledge graph update.
[0066] According to embodiments of this application, after acquiring and verifying adversarial examples, entity extraction of the adversarial examples is also performed. For example, after parsing the code data using an AST, entities (such as sensitive functions) are extracted and inserted into a preset knowledge graph. The construction of the knowledge graph will be described in detail below.
[0067] Figure 4 A flowchart illustrating the construction of a knowledge graph according to an embodiment of this application is shown.
[0068] like Figure 4As shown, according to an embodiment of this application, the knowledge graph is obtained by pre-setting, and the pre-setting process specifically includes operations S410 to S450.
[0069] When operating the S410, acquire heterogeneous security data from multiple sources, including vulnerability reports.
[0070] When operating the S420, entities and relationships in heterogeneous security data are extracted, and triples are stored. Triples include attack surface, vulnerability type, and affected components.
[0071] For example, multi-source heterogeneous security data can be aggregated, such as vulnerability reports (unstructured text), static code analysis results (code tree AST / CFG data stream), dynamic penetration test logs (penetration test tool output), and threat intelligence (threat intelligence engine, etc.).
[0072] Furthermore, natural language processing and data mining techniques are used to identify key entities and their relationships, such as attack surfaces (e.g., applications, operating systems), vulnerability types (e.g., buffer overflows), and affected components (e.g., databases, user interfaces). The relationships between entities are then determined, such as how a particular vulnerability type affects a specific component. Finally, the extracted entities and relationships are stored as triples in the format (attack surface, vulnerability type, affected component) and written to a knowledge base.
[0073] When operating S430, based on entities and relationships, taint tracing is used to generate causal relationship chains.
[0074] By operating the S440, a vulnerability attack is executed based on the causal chain, and the attack results are obtained.
[0075] When operating the S450, based on the verification of the attack results, a knowledge graph is constructed after the attack chain is determined.
[0076] For example, based on the extracted entities and relationships, taint tracking is performed, and causal relationship chains are generated through set rules and algorithms. In this process, the system analyzes attack paths, identifies possible attack flows, and records key nodes in these flows to form a complete causal chain.
[0077] Furthermore, after confirming the causal chain, simulation tools can be used to test the identified attack paths for vulnerability exploitation. Through penetration testing or other means, the system attempts to exploit the identified vulnerabilities and records the attack results, including success or failure, the scope of the attack, and the components attacked.
[0078] According to embodiments of this application, by integrating static taint tracking and dynamic penetration testing results, the attack chain detection coverage is improved.
[0079] Finally, based on the obtained attack results, the system will verify the validity of the attack chain. Based on the determined attack chain, a corresponding knowledge graph will be constructed. This knowledge graph will include the following: detailed information on various entities, relationships between entities, attack paths and corresponding vulnerability information, and analysis of attack results and impacts.
[0080] It should be noted that, in order to further optimize the graph structure of the constructed knowledge graph, the embodiments of this application define heterogeneous edge types of heterogeneous security data and dynamic weight calculation to achieve deep fusion of multi-source data; wherein, the heterogeneous edge types include static data streams, dynamic attack chains and text inference relationships.
[0081] It is understandable that different types of relationships (edges) connect entities (nodes) in a knowledge graph. Therefore, in the embodiments of this application, various edge types can be defined to reflect different information flows and interaction methods, based on the specific data source and the nature of the relationships, in order to better integrate security data from different sources.
[0082] Furthermore, when constructing a knowledge graph, by flexibly calculating weights, the strength of relationships between entities is no longer determined by fixed weight values, but dynamically adjusted based on factors such as context, time, and event frequency, which enables more efficient information integration and knowledge extraction.
[0083] According to an embodiment of this application, after the knowledge graph is constructed, multi-hop reasoning is implemented based on a relational graph convolutional network to predict potential vulnerability exploitation paths.
[0084] Specifically, a relational graph convolutional network is applied to train the knowledge graph, taking the feature vectors of input nodes (e.g., the severity of vulnerabilities, the frequency of component usage, etc.) and information about their neighboring nodes. Then, multiple rounds of convolutional operations are performed to aggregate and update the node features, thereby capturing more complex patterns and relationships.
[0085] Furthermore, multi-hop reasoning is achieved through a relational graph convolutional network model, allowing the system to determine possible exploitation paths starting from a known vulnerability. For example, starting from "vulnerability A", the system can reason about the path through "component B" and "component C" to finally reach the "target system".
[0086] Understandably, after generating potential exploit paths through the above steps, the inferred potential exploit paths can be fed back, which not only improves the understanding of existing vulnerabilities but also helps predict possible future attack paths.
[0087] Based on the above embodiments, the knowledge graph construction was completed. The following section explains how adversarial examples are parsed and the extracted entities are input into a preset knowledge graph to update the knowledge graph. It should be noted that the following detailed description of the embodiments uses the modality of the adversarial examples as code.
[0088] Figure 5 A flowchart illustrating the input of extracted entities into a knowledge graph according to an embodiment of this application is shown.
[0089] like Figure 5 As shown, according to an embodiment of this application, adversarial examples are parsed and the extracted entities are input into a preset knowledge graph, specifically including operations S510 to S530.
[0090] When operating the S510, the adversarial examples are parsed using a pre-defined parsing tool to construct an abstract syntax tree structure.
[0091] In the S520 operation, the abstract syntax tree structure is traversed to extract entities.
[0092] When operating S530, entities are input into a preset knowledge graph.
[0093] For example, parsing tools are used to analyze the format, features, and information of adversarial examples and construct the corresponding AST (Abstract Syntax Tree). During the AST construction process, the tool extracts key elements from the adversarial example, such as input features, perturbation methods, and attack targets.
[0094] Furthermore, the constructed abstract syntax tree is traversed, each node and its relationships are analyzed, and then the extracted entities are input into the knowledge graph one by one to form new nodes. If an extracted entity is related to an existing node in the knowledge graph, corresponding edges are added to these nodes to reflect the relationship between them. For example, "adversarial example - attack method" can form a new edge, linking the attack method to the specific adversarial example. Finally, after the entity insertion is completed, the knowledge graph is updated to support subsequent queries and enhanced retrieval.
[0095] When operating the S240, multimodal retrieval is performed based on knowledge graph updates to detect vulnerability information in adversarial samples.
[0096] According to embodiments of this application, multimodal retrieval is performed based on knowledge graph updates to detect vulnerability information in adversarial examples, including: determining vulnerability information in adversarial examples through adversarial training and counterfactual reasoning of three-level retrieval; wherein, the three-level retrieval includes sparse retrieval, dense retrieval, and reordering retrieval.
[0097] For example, this three-level retrieval includes, firstly, sparse retrieval, which is trained by utilizing key entities and attributes in a knowledge graph and performing sparse retrieval through keyword matching techniques. For instance, searching for "attack methods" and "vulnerabilities" related to "adversarial examples" to identify vulnerability information relevant to the current adversarial example.
[0098] Furthermore, intensive retrieval is performed. This intensive retrieval training involves analyzing similarities using a deep learning model based on the results of sparse retrieval, and then performing intensive retrieval. Potential vulnerabilities are uncovered by calculating the feature vector similarity between adversarial examples and existing adversarial examples in the knowledge graph.
[0099] Furthermore, a reordering retrieval process is performed. This reordering retrieval training involves applying a reordering algorithm to optimize the output order of the results of dense retrievals, and then ensuring that the most valuable vulnerability information is obtained based on relevance and importance.
[0100] In the embodiments of this application, after obtaining preliminary vulnerability information, counterfactual reasoning is also performed, that is, considering "what impact would changes in certain features of the adversarial example have on the model's prediction results." Further simulations are conducted to determine whether the model's output changes under different feature combinations, thereby identifying potential vulnerabilities. For example, after changing certain input features of the adversarial example, the changes in the model's output are observed; if a significant difference appears in the output, it may indicate that the feature is related to a vulnerability.
[0101] Finally, the vulnerability information obtained through multimodal retrieval and counterfactual reasoning will be integrated to form a final report, including the vulnerability description, characteristics, and possible attack paths.
[0102] According to embodiments of this application, based on the detected vulnerability information, the noise data used by the generator when constructing adversarial examples will also be adjusted to complete the generator's relearning. It is understood that the detected vulnerability information is fed back to the adversarial example generation stage, which allows for adjustments to the noise injection strategy, enabling the entire vulnerability detection model to form a dynamic iterative closed loop and continuously optimize the generation and detection effects.
[0103] Through the embodiments of this application, a multimodal retrieval strategy combining sparse retrieval, dense retrieval, and reordering retrieval enables the identification and extraction of vulnerability information in adversarial examples with higher efficiency. Furthermore, the application of counterfactual reasoning allows for more in-depth vulnerability analysis, simulating the possible behavior of models under different conditions to uncover potential vulnerabilities.
[0104] Based on the above vulnerability detection method, this application also provides a vulnerability detection device. The following will combine... Figure 6 The device is described in detail.
[0105] Figure 6 A schematic block diagram of a vulnerability detection device according to an embodiment of this application is shown.
[0106] like Figure 6 As shown, the vulnerability detection device 600 in this embodiment includes an original sample acquisition module 610, an adversarial sample generation module 620, a knowledge graph update module 630, and a vulnerability information detection module 640.
[0107] The original sample acquisition module 610 is used to acquire original samples, which are multimodal samples. In one embodiment, the original sample acquisition module 610 can be used to perform the operation S210 described above, which will not be repeated here.
[0108] The adversarial example generation module 620 is used to input multimodal samples into a preset generator and perform adversarial verification after obtaining the adversarial samples. In one embodiment, the adversarial example generation module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0109] The knowledge graph update module 630 is used to parse adversarial examples and input the extracted entities into a preset knowledge graph to complete the knowledge graph update. In one embodiment, the knowledge graph update module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0110] The vulnerability information detection module 640 is used to perform multimodal retrieval based on knowledge graph updates to detect vulnerability information in adversarial samples. In one embodiment, the vulnerability information detection module 640 can be used to perform the operation S240 described above, which will not be repeated here.
[0111] In the embodiments of this application, the adversarial sample generation module 620 includes an adversarial sample generation submodule and an adversarial verification submodule.
[0112] The adversarial example generation submodule is used to input multimodal samples into a preset generator, extract modal features through a dual-path coding structure, and perform feature fusion to obtain adversarial examples. In one embodiment, the adversarial example generation submodule can be used to perform the operation S310 described above, which will not be repeated here.
[0113] The adversarial verification submodule is used to input adversarial samples into a preset discriminator, perform multi-dimensional verification on the adversarial samples to obtain adversarial loss, and then send it back to the generator. In one embodiment, the adversarial verification submodule can be used to perform the operation S320 described above, which will not be repeated here.
[0114] In the embodiments of this application, the adversarial sample generation submodule includes a feature processing unit, which is included in a preset generator to perform feature alignment and cross-modal fusion processing on multimodal samples.
[0115] In the embodiments of this application, the knowledge graph update module 630 includes a knowledge graph construction submodule and an entity input submodule.
[0116] The knowledge graph construction submodule includes a data acquisition unit, an entity extraction unit, a relationship chain generation unit, an attack result acquisition unit, and a knowledge graph construction unit.
[0117] The data acquisition unit is used to acquire heterogeneous security data from multiple sources, including vulnerability reports. In one embodiment, the data acquisition unit can be used to perform the operation S410 described above, which will not be repeated here.
[0118] The entity extraction unit is used to extract entities and relationships from heterogeneous security data and store triples, which include the attack surface, vulnerability type, and affected components. In one embodiment, the entity extraction unit can be used to perform the operation S420 described above, which will not be repeated here.
[0119] The relationship chain generation unit is used to generate a causal relationship chain based on entities and relationships through taint tracking. In one embodiment, the relationship chain generation unit can be used to perform the operation S430 described above, which will not be repeated here.
[0120] The attack result acquisition unit is used to execute a vulnerability attack based on a causal chain and obtain the attack result. In one embodiment, the attack result acquisition unit can be used to execute the operation S440 described above, which will not be repeated here.
[0121] The knowledge graph construction unit is used for verification based on attack results. After determining the attack chain, it constructs a knowledge graph. In one embodiment, the knowledge graph construction unit can be used to perform the operation S450 described above, which will not be repeated here.
[0122] In the embodiments of this application, the entity input submodule includes a syntax tree construction unit, an entity extraction unit, and an entity input unit.
[0123] The syntax tree construction unit is used to parse adversarial examples based on a preset parsing tool and construct an abstract syntax tree structure. In one embodiment, the syntax tree construction unit can be used to perform the operation S510 described above, which will not be repeated here.
[0124] The entity extraction unit is used to traverse the abstract syntax tree structure and extract entities. In one embodiment, the entity extraction unit can be used to perform the operation S520 described above, which will not be repeated here.
[0125] The entity input unit is used to input entities into a preset knowledge graph. In one embodiment, the entity input unit can be used to perform the operation S530 described above, which will not be repeated here.
[0126] According to embodiments of this application, any multiple modules among the original sample acquisition module 610, adversarial sample generation module 620, knowledge graph update module 630, and vulnerability information detection module 640 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the original sample acquisition module 610, adversarial sample generation module 620, knowledge graph update module 630, and vulnerability information detection module 640 can be at least partially implemented as hardware circuitry, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the original sample acquisition module 610, adversarial sample generation module 620, knowledge graph update module 630, and vulnerability information detection module 640 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0127] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a vulnerability detection method according to an embodiment of this application.
[0128] like Figure 7 As shown, an electronic device 700 according to an embodiment of this application includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0129] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.
[0130] According to embodiments of this application, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0131] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0132] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used or combined with an instruction-modified execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.
[0133] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the vulnerability detection method provided in the embodiments of this application.
[0134] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0135] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0136] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0137] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0139] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0140] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A vulnerability detection method, characterized in that, The method includes: Obtain the original samples, which are multimodal samples; The multimodal samples are input into a preset generator, and adversarial verification is completed after adversarial samples are obtained. The adversarial examples are parsed, and the extracted entities are input into a preset knowledge graph to update the knowledge graph. Based on the updated knowledge graph, multimodal retrieval is performed to detect vulnerability information in the adversarial sample; In the preset generator, feature alignment and cross-modal fusion are performed on the multimodal samples.
2. The vulnerability detection method according to claim 1, characterized in that, The method further includes: Based on the vulnerability information, the noise data used by the generator when constructing the adversarial sample is adjusted to complete the generator's relearning.
3. The vulnerability detection method according to claim 1, characterized in that, The step of inputting the multimodal samples into a preset generator and completing adversarial verification after obtaining adversarial samples includes: The multimodal samples are input into a preset generator, and modal features are extracted and fused through a dual-path coding structure to obtain adversarial samples. The adversarial sample is input into a preset discriminator, and the adversarial loss is obtained by performing multi-dimensional verification on the adversarial sample and then sent back to the generator.
4. The vulnerability detection method according to claim 3, characterized in that, The step of inputting the multimodal samples into a preset generator and completing adversarial verification after obtaining adversarial samples further includes: Based on the multi-dimensional verification structure of the discriminator, the noise injection strategy and routing path selection of the generator are adjusted.
5. The vulnerability detection method according to claim 1, characterized in that, The knowledge graph is obtained through a pre-set process, which includes: Acquire heterogeneous security data from multiple sources, including vulnerability reports; Extract entities and relationships from the heterogeneous security data and store triples, wherein the triples include attack surface, vulnerability type and affected components; Based on the entities and relationships, a causal relationship chain is generated through taint tracking; Based on the aforementioned causal chain, execute the vulnerability attack and obtain the attack results; Based on the verification of the attack results, the knowledge graph is constructed after the attack chain is determined.
6. The vulnerability detection method according to claim 5, characterized in that, The pre-setting process of the knowledge graph also includes: By defining the heterogeneous edge types of the heterogeneous security data and calculating dynamic weights, deep fusion of multi-source data is achieved. The heterogeneous edge types include static data streams, dynamic attack chains, and text derivation relationships.
7. The vulnerability detection method according to claim 5, characterized in that, The pre-setting process of the knowledge graph also includes: Once the knowledge graph is constructed, multi-hop reasoning is achieved based on the relational graph convolutional network to predict potential vulnerability exploitation paths.
8. The vulnerability detection method according to claim 1, characterized in that, The adversarial examples are parsed, and the extracted entities are input into a preset knowledge graph, including: The adversarial sample is parsed using a pre-defined parsing tool to construct an abstract syntax tree structure; Traverse the abstract syntax tree structure and extract entities; The entity is input into a preset knowledge graph.
9. The vulnerability detection method according to claim 1, characterized in that, The process of performing multimodal retrieval based on the updated knowledge graph to detect vulnerability information in the adversarial sample includes: Vulnerability information in the adversarial samples is determined through adversarial training and counterfactual reasoning using a three-level retrieval system. The three-level retrieval system includes sparse retrieval, dense retrieval, and reordering retrieval.
10. A vulnerability detection device, characterized in that, The device includes: The original sample acquisition module is used to acquire original samples, which are multimodal samples; The adversarial sample generation module is used to input the multimodal samples into a preset generator and complete the adversarial verification after obtaining the adversarial samples. The knowledge graph update module is used to parse the adversarial sample and input the extracted entities into a preset knowledge graph to complete the update of the knowledge graph. The vulnerability information detection module is used to perform multimodal retrieval based on the update of the knowledge graph in order to detect vulnerability information in the adversarial sample; In the preset generator, the multimodal samples undergo feature alignment and cross-modal fusion processing.
11. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.