Multi-modal feature fusion software supply chain vulnerability intelligent positioning method
Through the software supply chain intelligent positioning method of multimodal feature fusion, federated learning, generative adversarial network, multi-head cross-attention mechanism, graph attention network and differentiable reinforcement learning framework are used to solve the problems of missing vulnerability representation dimensions and difficulty in fusion of cross-modal feature due to single modal feature analysis, and efficient vulnerability positioning and propagation path analysis are achieved, improving the comprehensiveness and accuracy of vulnerability detection in software supply chain.
Patent Information
- Application Number
- CN202510549549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In the detection of vulnerability chain vulnerability, the existing technology has lost the vulnerability representation dimension due to single modal feature analysis. The semantic gap between static code structure, dynamic behavior and text semantic data has caused difficulty in fusion of cross-modal features, resulting in feature redundancy and noise interference; the long-tail distribution characteristics and high cost of manual labeling of high-risk vulnerabilities have led to the intensification of sample scarcity, and the model is prone to overfitting local noise features under small sample conditions; the grammatical differences in different programming languages and cross-project dependence ecological diversity to split the multimodal feature space, and the existing transfer learning technology is difficult to achieve effective decoupling of code grammar and vulnerability logic, which restricts the model's cross-domain generalization ability; the fusion of dependency graphs and multimodal data is insufficient, and the analysis of vulnerability propagation paths is limited by the local association of component call chains, and it is impossible to accurately track the vulnerability diffusion paths between supply chain levels, resulting in insufficient vulnerability positioning accuracy in complex supply chain environments.
The software supply chain vulnerability intelligent positioning method using multimodal feature fusion is used to encrypt and aggregate the code structure data, dependent component metadata and runtime behavior log of distributed nodes through the federated learning framework to generate an encrypted feature set across nodes; the encrypted feature set is input into the conditional form to generate an adversarial network, and a code variant consistent with the semantics of the real vulnerability scenes is generated, and the similarity of the dynamic behavior sequence is compared through the discriminator network to output multimodal training samples; the cross-modal between code structure, text semantics and dynamic behavior is extracted using the multi-head cross-attention mechanism. The fusion feature is generated by the dependent component metadata, and the supply chain dependency graph is constructed through the graph attention network based on the dependent component metadata, quantify the probability of vulnerability propagation between components, and input the fusion feature vector into the graph attention network to identify the high-risk dependency paths associated with the code structure nodes in the dependent graph; adopt a differentiable reinforcement learning framework, and based on the vulnerability propagation probability of high-risk dependency paths and the attention weight of cross-modal correlation features, the target code is serialized and scanned, dynamically adjust the vulnerability positioning threshold strategy, and output the vulnerability code segment and the associated dependency chain.
Significantly improve the comprehensiveness and accuracy of vulnerability detection in software supply chains. Through multimodal feature fusion and collaborative optimization mechanisms, it alleviates the problem of model overfitting in small sample scenarios, eliminates the semantic gap between code, text and behavioral data, reduces noise interference, achieves accurate adaptation of risk scores and scanning ranges, improves cross-ecological generalization capabilities, accurately tracks the vulnerability propagation paths of complex dependency chains, and provides efficient and reliable technical support for the security governance of software supply chains.
Smart Images

Figure CN120068095A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an intelligent positioning method for software supply chain vulnerabilities with multi-modal feature fusion. Background Art
[0002] Software supply chain vulnerability detection refers to analyzing third-party libraries, toolchains, and dependent components introduced during the software development process to identify security risks caused by external code defects or malicious implants. Software supply chain vulnerability detection often relies on single-dimensional feature analysis, such as statically detecting code syntax structures or dynamically monitoring runtime abnormal behaviors, and fails to effectively integrate multi-source heterogeneous data, resulting in insufficient representation capabilities for complex vulnerability patterns. Due to the long-tail distribution characteristics of high-risk vulnerabilities and the annotation dependence on security expert experience, bottlenecks such as sample scarcity and poor cross-project generalization ability are faced in actual scenarios.
[0003] In the field of software supply chain vulnerability detection, existing technologies usually rely on single-modal feature analysis, such as static code structures or dynamic runtime behaviors, and it is difficult to comprehensively capture the multi-dimensional features of vulnerabilities. Due to the long-tail distribution characteristics of vulnerabilities and sample scarcity, the annotation of high-risk vulnerabilities highly depends on manual verification by security experts, resulting in a limited scale of the annotation dataset and a single coverage scenario. In addition, there is a semantic gap between multi-modal data (such as code structures, text semantics, and dynamic behaviors), and existing methods face problems of model parameter inflation and noise overfitting when performing cross-modal feature alignment and fusion. In cross-project or cross-language scenarios, differences in coding styles, dependency ecosystems, and vulnerability patterns among different code libraries further fragment the feature space, and existing transfer learning technologies are difficult to effectively decouple the complex association between code syntax and vulnerability logic, resulting in insufficient model generalization ability. Traditional methods do not fully integrate the supply chain dependency graph and multi-modal data, and the accuracy of vulnerability propagation path tracing is limited, making it difficult to cope with the complex security risks in the continuously evolving software supply chain environment. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technologies, the present invention provides an intelligent software supply chain vulnerability localization method based on multi-modal feature fusion, which solves the problems in the existing technologies. In the existing software supply chain vulnerability detection methods, due to single-modal feature analysis, the lack of vulnerability characterization dimensions occurs. The semantic gap between static code structures, dynamic behaviors, and text semantic data leads to difficulties in cross-modal feature fusion, resulting in feature redundancy and noise interference. The long-tailed distribution characteristics of high-risk vulnerabilities and the high cost of manual annotation exacerbate the scarcity of samples. Under the condition of small samples, the model is prone to overfitting local noise features. The syntax differences of different programming languages and the ecological diversity of cross-project dependencies fragment the multi-modal feature space. The existing transfer learning technologies are difficult to effectively decouple code syntax and vulnerability logic, restricting the cross-domain generalization ability of the model. The fusion of the dependency graph and multi-modal data is insufficient. The analysis of vulnerability propagation paths is limited to the local association of the component call chain and cannot accurately track the vulnerability diffusion paths between supply chain levels, resulting in insufficient vulnerability localization accuracy in complex software supply chain environments.
[0005] To solve the above technical problems, the specific technical solutions of the present invention are as follows: The intelligent software supply chain vulnerability localization method based on multi-modal feature fusion provided by the present invention includes: Performing multi-modal data collection on the target software system, and encrypting and aggregating the code structure data, dependent component metadata, and runtime behavior logs of distributed nodes through a federated learning framework to generate a cross-node encrypted feature set; Inputting the encrypted feature set into a conditional generative adversarial network, generating code variants with semantics consistent with real vulnerability scenarios through a generator network, and comparing the similarity of dynamic behavior sequences through a discriminator network to output multi-modal training samples; Inputting the multi-modal training samples into a multi-head cross-attention mechanism, extracting cross-modal correlation features between function call nodes, text entity embedding vectors, and dynamic event sequence encodings in the code structure diagram to generate a fused feature vector; Based on the dependent component metadata, constructing a supply chain dependency graph through a graph attention network, quantifying the vulnerability propagation probability between components, and inputting the fused feature vector into the graph attention network to identify high-risk dependency paths associated with code structure nodes in the dependency graph; Adopting a differentiable reinforcement learning framework, serially scanning the target code based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of the cross-modal correlation features, dynamically adjusting the vulnerability localization threshold strategy, and outputting vulnerable code segments and associated dependency chains.
[0006] Furthermore, in the intelligent software supply chain vulnerability localization method based on multi-modal feature fusion of the present invention, the federated learning framework includes: encrypting and aggregating the code structure data, dependent component metadata, and runtime behavior logs through a secure multi-party computation protocol to generate global feature statistics; Analyze the change frequency of the code structure data and the risk level of the dependent component metadata based on the reinforcement learning agent, dynamically adjust the data sampling weights of each distributed node, and preferentially collect the code features of high-frequency change modules and the version metadata of high-risk dependent components.
[0007] Furthermore, for the intelligent software supply chain vulnerability location method with multi-modal feature fusion according to the present invention, the conditional generative adversarial network includes: inputting the global feature statistics generated by the federated learning framework into the generator network, and constraining the generated code variants to conform to the abstract syntax tree rules through a code structure parser; Input the generated code variants and the dynamic behavior sequence into the discriminator network, and iteratively optimize the generator network to generate multi-modal training samples with consistent semantics.
[0008] Furthermore, for the intelligent software supply chain vulnerability location method with multi-modal feature fusion according to the present invention, the multi-head cross-attention mechanism includes: based on the fused feature vector, parallelly calculate the cross-modal association weights between the function call nodes in the code structure graph and the component nodes in the supply chain dependency graph; Perform hierarchical attention interaction on the embedding vector of the code structure node, the text entity description encoding, and the dynamic event sequence feature to generate a cross-modal consistent feature representation.
[0009] Furthermore, for the intelligent software supply chain vulnerability location method with multi-modal feature fusion according to the present invention, the graph attention network includes: updating the node risk features of the supply chain dependency graph based on the version information of the dependent component metadata and the vulnerability disclosure record; calculating the reachability probability of the vulnerability in the multi-layer dependency chain through the attention weight propagation algorithm, and using the reachability probability as the input parameter of the positioning strategy of the differentiable reinforcement learning framework.
[0010] Furthermore, for the intelligent software supply chain vulnerability location method with multi-modal feature fusion according to the present invention, the differentiable reinforcement learning framework includes: constructing a state space including code context semantic vectors, historical false alarm statistics, and dependency risk scores based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of the cross-modal association features; Set a reward function to evaluate the balance between the vulnerability location accuracy and the false alarm suppression effect, and optimize the positioning threshold decision parameters of the state space through the policy gradient algorithm.
[0011] Furthermore, the intelligent software supply chain vulnerability location method with multi-modal feature fusion according to the present invention further includes: Use the generated global feature statistics as the input of the generator of the conditional generative adversarial network to constrain the semantic consistency between the generated code variants and the global feature statistics; compare the dynamic behavior sequences of the generated adversarial samples and the real vulnerability samples through the discriminator network, and transmit the discrimination results back to the federated learning nodes to optimize the code structure feature extraction rules of the distributed nodes.
[0012] Further, the intelligent software supply chain vulnerability localization method with multi-modal feature fusion according to the present invention further includes: Use the component risk probability calculated by the graph attention network as the weight constraint condition of the multi-head cross-attention mechanism to suppress the correlation feature interaction between the code structure nodes and the low-risk dependency paths; Use the cross-modal consistency feature representation generated by the multi-head cross-attention mechanism to correct the edge weight parameters of the supply chain dependency graph and optimize the quantization accuracy of the vulnerability propagation probability.
[0013] Further, the intelligent software supply chain vulnerability localization method with multi-modal feature fusion according to the present invention further includes: Use the vulnerability propagation probability of the identified high-risk dependency paths as the prior knowledge of the state space of the differentiable reinforcement learning framework to dynamically narrow the scanning range of the target code; Use the vulnerability localization result output by the differentiable reinforcement learning framework to update the node risk features of the graph attention network in reverse.
[0014] Further, the intelligent software supply chain vulnerability localization method with multi-modal feature fusion according to the present invention further includes: Use a prototype network to cluster the vulnerability features of multi-modal training samples and construct a vulnerability pattern prototype vector for cross-programming language projects; Use the attention distillation technology to transfer the localization strategy parameters of the differentiable reinforcement learning framework to a lightweight interpretation model to generate heatmap interpretation information corresponding to the vulnerability code segments and associated dependency chains.
[0015] Advantages of the present invention; Through a multi-modal feature fusion and collaborative optimization mechanism, the present invention significantly improves the comprehensiveness and accuracy of software supply chain vulnerability detection. Based on the federated learning framework, it realizes the secure aggregation of cross-node data, combines generative adversarial networks to enhance sample diversity, and effectively alleviates the problem of model overfitting in small-sample scenarios; the multi-head cross-attention mechanism aligns cross-modal features of code structure, text semantics, and dynamic behavior, eliminates semantic gaps and reduces noise interference, and the graph attention network quantifies the vulnerability propagation probability of the dependency graph, jointly with the differentiable reinforcement learning framework to dynamically optimize the localization strategy, achieving precise adaptation of risk scores and scanning ranges; the prototype network and attention distillation technology construct cross-language vulnerability pattern prototypes, enhance model interpretability through heatmap explanations, and assist expert verification and closed-loop optimization. The synergistic effect of the above technologies breaks through the limitations of traditional single-modal analysis, improves cross-ecosystem generalization ability while ensuring data security, accurately tracks the vulnerability propagation path of complex dependency chains, and provides an efficient and reliable technical support for software supply chain security governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the drawings.
[0017] Figure 1 It is a flowchart of the intelligent software supply chain vulnerability localization method with multi-modal feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the specific embodiments and corresponding drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. The following will, in conjunction with the drawings, detail the technical solutions provided by each embodiment of the present invention. To better understand the objectives of the present invention, the following will further describe the present invention in detail.
[0019] Please refer to Figure 1 , the intelligent software supply chain vulnerability localization method with multi-modal feature fusion provided by the present invention includes: Step S101, perform multi-modal data collection on the target software system, and encrypt and aggregate the code structure data, dependent component metadata, and runtime behavior logs of distributed nodes through the federated learning framework to generate a cross-node encrypted feature set; A technical closed-loop is formed through multi-modal data collection, federated learning encrypted aggregation, and dynamic sampling optimization. The collaborative collection of multi-source heterogeneous data provides multi-dimensional inputs for vulnerability detection. The data security guarantee mechanism of federated learning solves the problem of data sharing in distributed scenarios. The dynamic sampling strategy improves the representativeness of the feature set and the ability to focus on risks. The encrypted feature set, as the basis for subsequent generative adversarial networks and cross-modal alignment, its generation logic and data quality directly affect the characterization accuracy of the model for complex vulnerability patterns, thus supporting the effective implementation of the overall technical solution.
[0020] When performing multi-modal data collection on the target software system, the multi-modal data includes the following three categories: Code structure data: Parse the source code or binary file through a static analysis tool, and extract syntax features such as function call relationships, control flow graphs, and variable dependency relationships in the abstract syntax tree. Such data represents the static logical structure of the code, such as function nesting levels, conditional branch paths, and data flow transfer rules, providing an analysis basis at the syntax level for vulnerability detection.
[0021] Dependency component metadata: The dependency management tool scans project configuration files (such as pom.xml or package.json) to obtain the version numbers, license types, component call levels, and compatibility declarations of third-party libraries. Such data reflects the external dependency ecosystem of the software supply chain, such as historical vulnerability records of high-risk component versions and the topological relationship of the dependency chain, and is used to evaluate the potential risks introduced by dependencies.
[0022] Runtime behavior logs: Monitor the execution process of the target system through dynamic probes, and collect memory access traces, system call sequences, inter-process communication events, and abnormal signal trigger timings. Such data records the actual behavior patterns during program execution, such as the characteristics of buffer overflow operations and the combined paths of abnormal system calls, providing behavioral evidence for the dynamic vulnerability trigger mechanism.
[0023] The three types of data respectively construct a heterogeneous feature space from three dimensions: code syntax, dependency ecosystem, and dynamic behavior, covering the static logic, supply chain context, and runtime semantics required for vulnerability characterization. The collaborative collection of multi-modal data provides multi-source inputs for subsequent federated learning encrypted aggregation, cross-modal feature alignment, and vulnerability propagation path analysis, supporting the full-life cycle detection of complex vulnerability patterns.
[0024] Step S102, input the encrypted feature set into a conditional generative adversarial network, generate code variants that are semantically consistent with real vulnerability scenarios through the generator network, and compare the similarity of dynamic behavior sequences through the discriminator network to output multi-modal training samples; When the encrypted feature set is input into the generator network, the encrypted features are first homomorphically decrypted and feature decoded to restore them to parsable code semantic features. The generator network has a built-in code structure parsing module that verifies the grammatical constraints of feature vectors based on abstract syntax tree rules to limit the variable scope closure, control flow integrity, and function call legality of the generated code. Through a multi-layer Transformer architecture, the cross-node vulnerability pattern features in the global feature statistics are fused with the local code syntax features to generate code variants that comply with programming language specifications and carry potential vulnerability semantics, such as function interfaces that do not verify input parameters or loop structures that have buffer overflow risks.
[0025] The discriminator network receives the code variants and real vulnerability samples output by the generator, and collects the runtime behavior sequences of the two in the simulated execution environment through dynamic probes. The behavior sequence includes memory access traces, system call chains, and abnormal signal triggering timings. After the multi-scale timing features are extracted by the temporal convolutional network, the similarity score between the generated samples and the real samples in the behavior pattern is calculated. The discriminator constructs an adversarial loss function, generates a gradient signal based on the difference in the behavior sequence, and back-propagates it to the generator network to adjust the parameters, driving the generation strategy to align with the semantic features of the real vulnerability scenario.
[0026] The adversarial training of the generator and the discriminator forms an iterative optimization process. The generator explores the feature combination of vulnerability patterns under grammatical constraints and generates code variants with potential threats; the discriminator selects high-authenticity samples based on the differences in dynamic behavior characteristics to form an adversarial feedback mechanism. After each round of training, the generated samples that meet the semantic consistency standards are retained and combined with the real samples to form a multimodal training sample set, covering the multi-dimensional features of code structure, dynamic behavior and vulnerability semantics. This sample set alleviates the risk of overfitting the model to local noise in small sample scenarios by expanding data diversity.
[0027] The generative adversarial network and the federated learning framework form a collaborative optimization mechanism. The behavioral difference features output by the discriminator are fed back to the federated learning node through the gradient return path to guide the local feature extractor to optimize the code structure parsing rules. For example, when the discriminator identifies that the generated sample has significant differences from the real sample in terms of memory out-of-bounds behavior, the federated learning node enhances the parsing weight of the memory operation mode and improves the characterization ability of subsequent encrypted aggregation features. In the process of data enhancement, the generative adversarial network inherits the global vulnerability pattern characteristics of federated learning, forms semantic consistency constraints across nodes, and enhances the coverage breadth of generated samples for multi-source vulnerability scenarios.
[0028] The above steps construct a closed-loop optimization link through encrypted feature decoding, adversarial sample generation, and dynamic behavior verification. The generator generates highly authentic code variants under the dual constraints of syntax rules and global features. The discriminator filters valid samples based on the similarity of behavior sequences. The multi-modal training sample set provides data support for subsequent cross-modal feature alignment. The collaborative feedback mechanism of the generative adversarial network and federated learning improves the semantic consistency of sample generation and the robustness of the model while protecting data security, laying a data foundation for the detection of complex vulnerability patterns.
[0029] Step S103: Input the multi-modal training samples into the multi-head cross-attention mechanism to extract the cross-modal correlation features between the function call nodes, text entity embedding vectors, and dynamic event sequence encodings in the code structure diagram, and generate a fused feature vector. When inputting the multi-modal training samples into the multi-head cross-attention mechanism, first perform graph embedding processing on the function call nodes in the code structure diagram and convert them into high-dimensional vector representations. The embedding vectors of the function call nodes extract hierarchical control flow features through a graph convolutional network, characterizing the call relationships between functions and the data transfer paths. The text entity embedding vectors use a pre-trained language model to perform semantic encoding on code comments, API documents, and vulnerability description texts to generate context semantic features associated with the code logic. The dynamic event sequence encoding extracts multi-scale temporal patterns of runtime behavior through a temporal convolutional network to capture the periodic patterns of abnormal memory operations or system calls.
[0030] The multi-head cross-attention mechanism constructs multiple independent attention subspaces in parallel, and each subspace calculates the interaction weights between different modal features. In the attention subspaces of code structure and dynamic behavior, the association strength between function call nodes and dynamic event sequences is quantified through vector dot product operations, such as identifying the temporal coupling relationship between a specific loop structure and a memory out-of-bounds operation. In the subspace of text semantics and code structure, the cosine similarity between the code node embedding vector and the text entity description vector is calculated to align the potential associations between code logic and document semantics. After the attention weight matrices of each subspace are normalized, they are concatenated to form a cross-modal correlation feature map, reflecting the multi-dimensional interaction patterns of code syntax, dynamic behavior, and text description.
[0031] The cross-modal associated feature map inputs a hierarchical attention interaction module for feature fusion. The first-layer attention aggregates the associated features of the code structure nodes and the dynamic event sequence, and filters out the code-behavior interaction patterns with high weight values, such as the associated path between a function call without verifying input parameters and subsequent buffer overflow events. The second-layer attention interacts the aggregation result with the text semantic features to enhance the semantic consistency between the code logic and the vulnerability description keywords, such as associating high-risk function names with risk keywords in the vulnerability database. During the hierarchical fusion process, residual connections are used to preserve the original modal features and avoid the loss of key information, and finally a fusion feature vector representing the multi-dimensional characteristics of the vulnerability is generated.
[0032] The fusion feature vector, as a unified multi-modal representation, is input into the subsequent graph attention network for node risk quantification of the supply chain dependency graph. The code-behavior interaction patterns extracted from the cross-modal associated features are used to correct the edge weight parameters of the dependency graph, and the text semantic alignment features assist in identifying potential risk keywords in the descriptions of the dependent components. The synergistic effect of the multi-head cross-attention mechanism and hierarchical fusion eliminates the modal gap between code syntax, dynamic behavior, and text semantics, provides a multi-dimensional feature input with consistent semantics for the analysis of vulnerability propagation in the dependency chain, and supports the accurate positioning of complex vulnerability patterns.
[0033] The above steps construct a cross-modal alignment link through graph embedding, multi-subspace attention calculation, and hierarchical feature fusion. The feature interaction mechanism of code structure, dynamic behavior, and text semantics breaks through the limitations of single-modal analysis; the generation logic of the fusion feature vector is connected with the subsequent graph analysis, forming a technical closed-loop from multi-modal alignment to dependency risk quantification, and improving the vulnerability detection model's perception ability for hidden vulnerabilities and adaptability to cross-ecosystem scenarios.
[0034] Step S104, based on the metadata of the dependent components, construct a supply chain dependency graph through a graph attention network, quantify the vulnerability propagation probability between components, and input the fusion feature vector into the graph attention network to identify high-risk dependency paths associated with code structure nodes in the dependency graph; When constructing a supply chain dependency graph based on the metadata of dependent components, first parse the component version information and vulnerability disclosure records, and map the version numbers, maintenance status, and historical vulnerability data of third-party libraries to the risk attributes of the graph nodes. The call relationships between components establish directed edges through the dependency declarations in the project configuration files, and the initial values of the edge weights are dynamically assigned according to the interface call frequency and component compatibility declarations. The graph attention network enhances the features of the node risk attributes and quantifies the inherent risk level in combination with the CVSS scores in the vulnerability database. For example, nodes of obsolete version components that have not been updated for a long time are given higher initial risk weights.
[0035] The graph attention network uses a hierarchical attention mechanism to calculate the vulnerability propagation probability between nodes. The first-layer attention aggregates the risk features of adjacent nodes and propagates the vulnerability influence range along the direction of the dependency chain through the message passing mechanism. The second-layer attention introduces a cross-modal fusion feature vector to associate the control flow features of code structure nodes with the risk attributes of dependent components, and calculates the semantic coupling strength between specific function call nodes and high-risk component nodes. After normalization, the attention weights generate the reachability probability of the multi-layer dependency chain, which characterizes the potential path of vulnerability diffusion from the underlying dependent components to the core code module.
[0036] After the fusion feature vector is input into the graph attention network, the cross-modal correlation features are mapped to the edge weight space of the dependency graph through the feature projection mechanism. The interaction features between code structure nodes and dynamic behavior sequences are used to correct the call frequency weights of dependency edges. For example, identify the interface dependency path with high-frequency calls but memory out-of-bounds risks. The correlation strength between text semantic features and vulnerability description keywords synchronously adjusts the node risk score, enhancing the ability to perceive implicit risks in the document. The dynamic adjustment of edge weights and node risks forms the topological optimization of the dependency graph, accurately quantifying the probability gradient of vulnerability propagation between components.
[0037] The identification of high-risk dependency paths is based on the joint screening of reachability probability thresholds and cross-modal attention weights. The graph attention network traverses all possible propagation paths in the dependency graph and marks the paths with reachability probabilities exceeding the preset threshold as high-risk candidates. At the same time, the semantic coupling strength between code nodes and dependent components in the fusion feature vector is used as an auxiliary criterion to screen out the dependency chains strongly associated with the core business logic. For example, identify the function call chain through the initialization function of a high-risk third-party library that affects the main program entry module, and verify its propagation feasibility by combining historical vulnerability data.
[0038] The above steps achieve the dynamic analysis of risk paths through dependency graph construction, propagation probability quantification, and feature vector fusion. The multi-level enhancement of node risk attributes and the projection correction of cross-modal features break through the static limitations of traditional dependency chain analysis; the joint screening mechanism of reachability probability and semantic coupling strength supports the accurate positioning of key vulnerability propagation paths in complex supply chain scenarios, providing a topological basis for the dynamic scanning strategy of subsequent reinforcement learning.
[0039] In step S105, a differentiable reinforcement learning framework is adopted to serially scan the target code based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of the cross-modal correlation features, dynamically adjust the vulnerability location threshold strategy, and output the vulnerable code segment and the associated dependency chain.
[0040] When adopting a differentiable reinforcement learning framework, first construct a state space based on the vulnerability propagation probability of high-risk dependency paths and the attention weights of cross-modal correlation features. The state space integrates code context semantic vectors, historical false alarm statistics, and dependency risk scores to form multi-dimensional input features. The code context semantic vectors are extracted through a multi-head cross-attention mechanism, representing the comprehensive features of the syntax structure, dynamic behavior patterns, and text description semantics of the target code segment; the historical false alarm statistics extract the distribution law of false alarm cases from system logs, quantifying the correlation between the false alarm frequencies of different code modules and detection rules; the dependency risk score is dynamically updated based on the reachability probability calculated by the graph attention network, reflecting the real-time evaluation results of the vulnerability propagation risk in the dependency chain.
[0041] When setting the reward function, adopt a multi-objective weighted strategy to balance the vulnerability location accuracy and false alarm suppression effect. The location accuracy calculates the gain value through the recall rate and precision rate of the labeled vulnerability samples, and the false alarm suppression effect generates a penalty term based on the historical false alarm statistics. The reward function takes the difference between the recall gain and false alarm penalty as the optimization direction, and adjusts the sensitivity threshold in different scenarios through a dynamic weight allocation mechanism. For example, when the dependency risk score is relatively high and the cross-modal attention weights are concentrated in a specific code area, increase the weight coefficient of the recall rate to enhance the scanning intensity of high-risk areas; when the historical false alarm data shows that a certain type of code pattern frequently triggers false alarms, increase the false alarm penalty weight of the corresponding area to suppress over-detection.
[0042] When optimizing the location threshold decision parameters of the state space through the policy gradient algorithm, map the multi-dimensional features of the state space to a set of candidate thresholds, and calculate the policy gradient based on the evaluation signal of the reward function. The location threshold parameter controls the scanning granularity priority of different code areas. For example, enable low-threshold and high-sensitivity detection for modules associated with high-risk dependency paths, and apply high thresholds to low-risk areas to reduce computational overhead. Monte Carlo sampling explores the decision effects of different threshold combinations, and the gradient ascent method updates the parameters, making the threshold strategy with high reward values obtain a higher selection probability, and gradually converging to the optimal parameter distribution.
[0043] When outputting the vulnerability location results, combine the dynamic scanning strategy to screen out high-risk code segments and associated dependency chains. During the serialized scanning process, first perform fine-grained detection on the code associated with the dependency paths whose reachability probability exceeds the preset threshold, and extract the code lines and call stack information suspected of having vulnerabilities. The associated dependency chain generates a complete call link from the underlying high-risk component to the target code module by backtracking the propagation path in the dependency graph. For example, mark the specific path that causes a memory leak in the main program through the initialization function of a third-party library.
[0044] The positioning results are fed back to the graph attention network through the gradient backpropagation mechanism to form a closed-loop optimization. For the code segments identified as vulnerabilities, the risk scores of the corresponding graph nodes are enhanced according to the actual propagation of their associated dependency paths; for false positive cases, the risk weights of the relevant nodes are adjusted downward in combination with the deviation degree between the code context semantics and the dependency chain. The updated node features are re-input into the graph attention network for edge weight correction, realizing the dynamic collaborative iteration of the dependency graph risk quantification model and the reinforcement learning positioning strategy, and enhancing the adaptive detection ability in complex supply chain environments.
[0045] The above steps form a technical closed-loop through state space construction, reward mechanism design, policy optimization, and closed-loop feedback. The multi-dimensional feature fusion provides a decision-making basis for reinforcement learning, the dynamic threshold adjustment balances the detection efficiency and accuracy, and the result feedback mechanism realizes the co-evolution of the dependency graph and the positioning strategy. The deep integration of the reinforcement learning framework with cross-modal features and the graph attention network breaks through the limitations of traditional static detection rules and enables the accurate tracking and real-time response of complex vulnerability propagation paths.
[0046] The intelligent software supply chain vulnerability positioning method with multi-modal feature fusion provided by the present invention realizes the complete logical closed-loop of the technical solution through the following steps: First, multi-modal data collection is performed on the target software system, and the encrypted aggregation of the code structure data, dependency component metadata, and runtime behavior logs of distributed nodes is carried out using the federated learning framework to generate a cross-node encrypted feature set. In this process, the federated learning framework encrypts the code syntax tree, dependency version information, and runtime memory operation logs uploaded by each node through a secure multi-party computation protocol, and aggregates them to generate global feature statistics, avoiding the risk of raw data leakage. At the same time, based on the reinforcement learning agent, the code structure change frequency and the risk level of dependency components are analyzed, and the data sampling weights of each node are dynamically adjusted to preferentially collect the code features of frequently modified modules and the metadata of high-risk dependency components with known vulnerabilities, improving the coverage and efficiency of data collection.
[0047] When the encrypted feature set is input into the conditional generative adversarial network for data augmentation, the generator network generates code variants that conform to the abstract syntax tree rules under the constraint of the code structure parser, while maintaining semantic consistency with real vulnerability scenarios. The discriminator network compares the dynamic behavior sequences of the generated code variants and real vulnerability samples, and evaluates the authenticity of the generated samples through the abnormal memory access patterns or system call chain similarities in the dynamic behavior logs. Through the adversarial training iteration optimization of the generator and the discriminator, multi-modal training samples including code variants and corresponding dynamic behavior sequences are output, expanding the training data diversity in small sample scenarios and enhancing the robustness of the model against code obfuscation and adversarial attacks.
[0048] The multi-modal training samples are then input into a multi-head cross-attention mechanism for cross-modal feature alignment. This mechanism parallelly calculates the correlation weights between function call nodes, text entity embedding vectors, and dynamic event sequence encodings in the code structure diagram, and extracts cross-modal consistency features. Specifically, the function call relationships in the code structure diagram are transformed into vector representations through graph embedding techniques, the text entity descriptions are encoded into semantic vectors by a pre-trained language model, and the dynamic event sequences extract temporal features through a temporal convolutional network. The multi-head attention mechanism aligns and fuses features of different modalities in multiple semantic spaces through a hierarchical interaction strategy, generating a fused feature vector representing the multi-dimensional characteristics of vulnerabilities and eliminating the semantic gap between code syntax and dynamic behavior.
[0049] When constructing a supply chain dependency graph based on the metadata of dependent components, the graph attention network dynamically integrates component version information and vulnerability disclosure records, and updates the risk features of the graph nodes. The reachability probability of vulnerabilities in the multi-layer dependency chain is quantified through the attention weight propagation algorithm to identify the critical propagation paths between components. During this process, the fused feature vector is input into the graph attention network to associate the code structure nodes with the high-risk component nodes in the dependency graph, such as identifying the code functions that call high-risk third-party libraries, so as to locate the critical paths in the dependency chain that may cause vulnerability diffusion.
[0050] When using a differentiable reinforcement learning framework for vulnerability localization, based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of cross-modal correlation features, a state space including code context semantics, historical false positive statistics, and dependency risk scores is constructed. A reward function is set to evaluate the balance between localization accuracy and false positive suppression effect, and the vulnerability detection threshold is dynamically adjusted through the policy gradient algorithm. For example, when the dependency path risk score is high and the cross-modal attention weights are concentrated in a specific code segment, the reinforcement learning agent preferentially scans this area, optimizes the localization strategy in combination with historical false positive data, and finally outputs the vulnerable code segment and its associated dependency chain, realizing an accurate mapping from global dependency analysis to local code localization.
[0051] The above steps form a closed loop through the collaborative mechanism of data flow and model training: federated learning ensures the secure aggregation of multi-source data, generative adversarial networks improve sample diversity, multi-head attention mechanisms achieve cross-modal feature alignment, graph attention networks quantify dependency risks, and reinforcement learning frameworks dynamically optimize localization strategies. The outputs of each step serve as the inputs for subsequent steps. For example, the encrypted feature set drives data augmentation, the fused feature vector supports dependency graph analysis, and the dependency risk score guides reinforcement learning decisions, ultimately achieving a comprehensive improvement in vulnerability localization accuracy and interpretability in complex supply chain environments.
[0052] Specifically, for the intelligent positioning method of software supply chain vulnerabilities with multi-modal feature fusion described in the present invention, the federated learning framework includes: encrypting and aggregating the code structure data, dependent component metadata, and runtime behavior logs through a secure multi-party computation protocol to generate global feature statistics; Based on the reinforcement learning agent, analyze the change frequency of the code structure data and the risk level of the dependent component metadata, dynamically adjust the data sampling weights of each distributed node, and preferentially collect the code features of high-frequency change modules and the version metadata of high-risk dependent components.
[0053] The operation process of the federated learning framework described in the present invention is realized through the following technical solutions: when encrypting and aggregating the code structure data, dependent component metadata, and runtime behavior logs of each distributed node through a secure multi-party computation protocol, a homomorphic encryption algorithm is used to encrypt and transmit the code syntax tree feature vector, dependent version hash value, and runtime memory access sequence, and perform a feature weighted average operation in ciphertext state on the central aggregation server to generate global feature statistics. This process replaces the original data exchange with encrypted intermediate parameters, prevents the leakage of sensitive information in the code repository, and at the same time retains the semantic consistency of cross-node features.
[0054] On the basis of generating global feature statistics, the reinforcement learning agent analyzes the module activity index according to the change frequency of the code structure data, and combines the evaluation results of the risk level of the dependent component metadata to construct a dynamic sampling strategy. Specifically, the code change frequency calculates the number of function-level modifications through the commit records of the version control system, and the dependency risk level quantifies the threat degree of the component based on the CVSS score of the vulnerability database. The reinforcement learning agent uses the module activity and threat degree as state inputs, and optimizes the node sampling weight allocation strategy through the Q-learning algorithm, so that the code features of high-frequency modified modules and the version metadata of high-risk dependent components obtain priority collection permissions, improving the coverage efficiency of data collection and the risk focusing ability.
[0055] In the above operations, the secure multi-party computation protocol and the reinforcement learning strategy form a collaborative mechanism: encryption aggregation ensures the security of multi-source data, and the global feature statistics reflect the distribution law of cross-node data; the reinforcement learning agent dynamically optimizes the sampling strategy based on the global features, guiding the subsequent data collection to focus on high-risk areas. Data collection and strategy optimization form a closed-loop feedback. When a newly disclosed dependent component vulnerability is detected, the reinforcement learning agent automatically increases the sampling weight of the associated node, triggering targeted data supplementary collection to achieve dynamic response to changes in supply chain risks.
[0056] Specifically, for the intelligent location method of software supply chain vulnerabilities with multi-modal feature fusion described in the present invention, the conditional generative adversarial network includes: inputting the global feature statistics generated by the federated learning framework into the generator network, and constraining the generated code variants to conform to the abstract syntax tree rules through a code structure parser; Inputting the generated code variants and the dynamic behavior sequence into the discriminator network, and iteratively optimizing the generator network to generate multi-modal training samples with consistent semantics.
[0057] The operation process of the conditional generative adversarial network described in the present invention is realized through the following technical solutions: when inputting the global feature statistics generated by the federated learning framework into the generator network, the generator network built-in code structure parser decodes and reconstructs the input encrypted feature set. The code structure parser parses the code syntax features based on the abstract syntax tree rules, and constrains the code variants output by the generator to conform to the basic syntax structure of the programming language, such as the integrity of function definitions, the legality of variable scopes, and the closure of control flows, to avoid generating invalid code fragments that cannot be compiled or executed. Under the syntax constraint, the generator network combines the cross-node code pattern features in the global feature statistics to generate code variants with potential vulnerability semantics, such as functions that do not verify input parameters or loop structures with memory out-of-bounds risks.
[0058] The generated code variants and the dynamic behavior sequence of real vulnerability samples are jointly input into the discriminator network for adversarial training. The dynamic behavior sequence collects the abnormal system call chain, memory access pattern, and inter-process communication features triggered by real vulnerabilities through runtime probes to form a multi-dimensional time series behavior vector. The discriminator network uses a temporal convolutional layer to extract the dynamic behavior time series features of the code variants and real samples, and calculates the difference scores in dimensions such as memory operation frequency, system call path similarity, and abnormal event density. Based on the difference scores, an adversarial loss function is constructed, and the parameters of the generator network are iteratively optimized through gradient backpropagation, forcing the generator to generate code variants with dynamic behavior features highly matching those of real vulnerability samples, and improving the semantic authenticity of the generated samples.
[0059] During the adversarial training process of the generator and the discriminator, the syntactic compliance and behavioral authenticity of the generated samples form a double constraint. The code structure parser ensures the basic syntax correctness of the generated code, avoiding the adversarial training from falling into local noise optimization; the dynamic behavior similarity constraint guides the generator to learn the key behavior patterns triggered by vulnerabilities, such as repeated out-of-bounds accesses to specific memory addresses or combined sequences of abnormal system calls. After each round of training iteration, the code variants output by the generator will be added as new samples to the multi-modal training set, and participate in subsequent feature fusion and model training together with real samples, enhancing the robustness of the model to code variants and adversarial attacks.
[0060] In the above operations, the global feature statistics of federated learning provide prior knowledge of cross-node code patterns for the generative adversarial network, guiding the generator to focus on the feature generation of high-frequency vulnerability patterns; the discriminator establishes a consistency evaluation criterion for vulnerability semantics through dynamic behavior comparison, forming a closed-loop optimization mechanism for data augmentation. The synergistic effect of the generative adversarial network and federated learning is reflected in that federated learning ensures feature sharing under the security of multi-source data, and the generative adversarial network uses the shared features to generate diverse training samples, and the two jointly solve the problems of model overfitting and insufficient generalization ability in the small-sample scenario.
[0061] Specifically, for the intelligent software supply chain vulnerability location method with multi-modal feature fusion described in the present invention, the multi-head cross-attention mechanism includes: based on the fused feature vector, calculating the cross-modal association weights of function call nodes in the code structure graph and component nodes in the supply chain dependency graph in parallel; Performing hierarchical attention interaction on the embedding vector of the code structure node, the text entity description encoding, and the dynamic event sequence feature to generate a cross-modal consistency feature representation.
[0062] The operation process of the multi-head cross-attention mechanism of the present invention is realized through the following technical solutions: when calculating the cross-modal association weights of function call nodes in the code structure graph and component nodes in the supply chain dependency graph based on the generated fused feature vector, the graph embedding technology is used to map the function call relationship in the code structure graph into a vector space representation, and at the same time, the risk propagation features of components in the dependency graph are extracted through the graph attention network. The cross-modal association degree between the function call node vector and the dependent component node vector is calculated through vector dot product and cosine similarity, and an attention weight matrix reflecting the coupling degree of code logic and dependency risk is generated. This weight matrix is used to identify the potential impact area of high-risk dependency paths on the code structure, such as the entry function of a high-risk third-party library call or an interface module with vulnerability propagation risk.
[0063] When performing hierarchical attention interaction on the embedding vector of the code structure node, the text entity description encoding, and the dynamic event sequence feature, a hierarchical multi-head attention subspace is designed. The embedding vector of the code structure node extracts function-level control flow features through the graph convolutional network, the text entity description encoding is converted into a semantic vector through the pre-trained language model, and the dynamic event sequence feature uses a time sliding window to capture the abnormal behavior time series pattern. Each attention head independently calculates the interaction weights of different modal features in the subspace, and aggregates the multi-level feature interaction results through splicing and normalization operations. The hierarchical attention mechanism preferentially retains the association features between the code structure node and the high-risk dynamic event sequence, such as the time series correlation between memory out-of-bounds operations and specific function calls, and at the same time suppresses the interference of non-critical annotation information in the text description.
[0064] In the above operations, cross-modal correlation weight calculation and hierarchical attention interaction form a collaborative mechanism: the correlation weight matrix provides a cross-modal semantic alignment basis for hierarchical attention, guiding different attention heads to focus on the key interaction areas of code logic and dependency risks; the cross-modal consistency feature representation output by hierarchical interaction further optimizes the update direction of the correlation weight matrix. For example, it enhances the attention distribution intensity of high-risk dependent components to associated code nodes. The multi-modal features of code structure, dependency graph, and dynamic behavior are aligned and fused in the semantic space through this mechanism, providing a unified multi-dimensional feature input for subsequent vulnerability propagation path analysis and eliminating the locality and one-sidedness defects of single-modal analysis.
[0065] Specifically, for the software supply chain vulnerability intelligent location method with multi-modal feature fusion described in the present invention, the graph attention network includes: updating the node risk features of the supply chain dependency graph based on the version information and vulnerability disclosure records of the dependent component metadata; calculating the reachability probability of vulnerabilities in multi-layer dependency chains through the attention weight propagation algorithm, and using the reachability probability as the input parameter of the positioning strategy of the differentiable reinforcement learning framework.
[0066] The operation process of the graph attention network described in the present invention is realized through the following technical solutions: when updating the node risk features of the supply chain dependency graph based on the version information of the dependent component metadata and the vulnerability disclosure records, the release time, maintenance status, and compatibility statement in the component version information are parsed, and combined with the CVE number and CVSS score in the vulnerability disclosure records to quantify the inherent risk level of the component nodes. For example, abandoned version components that have not been updated for a long time are given a higher risk weight according to the vulnerability disclosure frequency, and the component versions with known high-risk vulnerabilities are superimposed with risk coefficients based on the CVSS score. The node risk features are mapped into multi-dimensional vectors through graph embedding technology to dynamically update the node attributes in the dependency graph, reflecting the real-time security situation of the supply chain components.
[0067] When calculating the reachability probability of vulnerabilities in multi-layer dependency chains through the attention weight propagation algorithm, the message passing method in the graph attention mechanism is adopted to aggregate the risk features of adjacent nodes along the edge connection direction of the dependency graph. For each node on the dependency path, calculate the attention weight between it and the upstream root node, and generate the path reachability probability after normalization through the softmax function. This probability represents the possibility of a vulnerability spreading from the underlying dependent components of the supply chain to the target code module. For example, the probability that a certain high-risk third-party library affects the main program entry function through a multi-level dependency call chain. During the calculation of the reachability probability, the edge weights of the dependency graph are dynamically adjusted according to the call frequency and interface complexity between components, and the dependency paths with high call frequency and high interface coupling degree are given priority attention.
[0068] When the reachability probability is used as an input parameter for the localization strategy of the differentiable reinforcement learning framework, the reinforcement learning agent weights and fuses the reachability probability with the attention weights of the cross-modal association features to construct a decision basis for vulnerability localization. For example, when the cross-modal attention weight of a certain code module is high and the reachability probability of the associated dependency path exceeds a preset threshold, the reinforcement learning agent preferentially performs a fine-grained scan on this module. During the localization process, the reachability probability affects the parameter update direction of the graph attention network through the gradient backpropagation mechanism. For example, when the localization result feedback shows that the actual vulnerability propagation risk of a certain dependency path is lower than the estimate, the reachability probability weight of the corresponding path is automatically adjusted downward to form a dynamic collaborative optimization of the dependency graph risk quantification and the localization strategy.
[0069] In the above operations, the node risk feature update and the reachability probability calculation form a hierarchical analysis structure: the inherent risk of the dependent component provides the basic attributes for the reachability propagation, and the graph attention mechanism quantifies the risk diffusion path through the topological structure analysis; the reinforcement learning framework fuses the reachability probability with multi-modal features to drive the dynamic optimization of the localization strategy. At the same time, the localization result reversely optimizes the risk features of the graph nodes, realizing a closed-loop feedback of dependent risk analysis - vulnerability localization - graph update. This mechanism effectively solves the problem of insufficient evaluation of the risk propagation of deep dependency chains in traditional methods and improves the accuracy of vulnerability tracing and localization in complex supply chain scenarios.
[0070] Specifically, for the intelligent software supply chain vulnerability localization method with multi-modal feature fusion described in the present invention, the differentiable reinforcement learning framework includes: constructing a state space including code context semantic vectors, historical false positive statistics, and dependency risk scores based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of the cross-modal association features; Setting a reward function to evaluate the balance between vulnerability localization accuracy and false positive suppression effect, and optimizing the localization threshold decision parameters of the state space through the policy gradient algorithm.
[0071] The operation process of the differentiable reinforcement learning framework described in the present invention is realized through the following technical solutions: when constructing the state space based on the vulnerability propagation probability of the high-risk dependency path and the attention weights of the cross-modal association features, the code context semantic vector is generated by the fused feature vector extracted by the multi-head cross-attention mechanism, representing the comprehensive features of the syntax structure, text description semantics, and dynamic behavior pattern of the target code segment. The historical false positive statistics are extracted through the false positive cases recorded in the system log to quantify the correlation between the false positive occurrence frequency of different code modules and the vulnerability detection rules. The dependency risk score is dynamically adjusted according to the reachability probability calculated by the graph attention network, reflecting the real-time evaluation result of the vulnerability propagation risk in the dependency path. The state space integrates the above multi-dimensional features to form an input benchmark for vulnerability localization decisions.
[0072] When setting the reward function to evaluate the balance between vulnerability location accuracy and false positive suppression effect, a weighted multi-objective optimization strategy is adopted. The location accuracy calculates the recall rate and precision rate through the detection results of the labeled vulnerability samples, and the false positive suppression effect generates a penalty term based on the pattern rules in the historical false positive statistics. The reward function takes the difference between the location accuracy gain and the false positive penalty term as the optimization direction, and balances the detection sensitivity and false positive control requirements through a dynamic weight allocation mechanism. For example, when the dependency risk score is high and the cross-modal attention weight is concentrated in a specific code area, the weight coefficient for improving the recall rate is increased to enhance the scanning intensity of high-risk areas; when the historical false positive statistics show that a certain type of code pattern frequently triggers false positives, the false positive penalty weight for the corresponding area is increased to suppress over-detection.
[0073] When optimizing the location threshold decision parameters in the state space through the policy gradient algorithm, the multi-dimensional features in the state space are mapped to a set of candidate location thresholds, and the policy gradient is calculated based on the evaluation signal output by the reward function. The location threshold decision parameters control the scanning granularity priority of different code areas. For example, for code modules associated with high-risk dependency paths, low thresholds and high sensitivity detections are used, and high thresholds are applied to low-risk areas to reduce the computational overhead. The policy gradient algorithm explores the decision effects of threshold combinations through Monte Carlo sampling, and updates the parameters using the gradient ascent method, so that the threshold strategy with high reward values has a higher selection probability, and gradually converges to the optimal decision parameter distribution.
[0074] In the above operations, the construction of the state space and the design of the reward function form the data basis for decision optimization: the multi-dimensional feature fusion reflects the complex working conditions of vulnerability detection, and the reward mechanism defines the directional constraints of the optimization goal. The policy gradient algorithm realizes the end-to-end association between the threshold parameters and the detection effect through a differentiable mechanism, so that the location strategy can adapt to code changes and dependency risk changes. The coordination of the reinforcement learning framework with the graph attention network and the multi-head cross-attention mechanism is reflected in: the dependency risk score and the cross-modal attention weight drive the dynamic reorganization of the features in the state space, and the location threshold decision result reversely triggers the update request of the features of the dependency graph nodes, forming a closed-loop iterative optimization of the detection strategy and risk perception, and improving the adaptive ability of vulnerability location in complex supply chain environments.
[0075] Specifically, the intelligent software supply chain vulnerability location method with multi-modal feature fusion described in the present invention further includes: using the generated global feature statistic as the input of the generator of the conditional generative adversarial network to constrain the semantic consistency between the generated code variant and the global feature statistic; comparing the dynamic behavior sequences of the generated adversarial samples and the real vulnerability samples through the discriminator network, and reversely transmitting the discrimination result to the federated learning nodes to optimize the code structure feature extraction rules of the distributed nodes.
[0076] The operation process of the cooperation between the generative adversarial network and federated learning in the present invention is realized through the following technical solutions: when the global feature statistics generated by the federated learning framework are input into the generator network, the generator network restores the encrypted aggregated statistics into parsable code semantic features through a feature decoder. The code structure parser performs syntax verification on the code variants output by the generator based on the abstract syntax tree rules, constraining the legality of variable scopes and the closure of control flows, so that the generated code variants inherit the cross-node vulnerability pattern features in the global feature statistics while maintaining syntax correctness. For example, the function call pattern with missing permissions that frequently appears in the global features will guide the generator to preferentially construct similar function structures in the code variants, and at the same time avoid generating invalid code through syntax verification.
[0077] When the generated code variants and the dynamic behavior sequences of real vulnerability samples are input into the discriminator network for adversarial training, the dynamic behavior sequences collect abnormal memory access traces and system call chains through runtime probes to form multi-dimensional time series feature vectors. The discriminator network uses temporal convolutional layers to extract the behavior pattern features during the execution of the code variants, and calculates the difference scores in dimensions such as the memory operation frequency, the interval between abnormal signal triggers, and the similarity of system call paths between it and the real samples. The difference scores are converted into gradient signals through an adversarial loss function to drive the generator network to adjust the semantic generation strategy of the code variants, gradually narrowing the behavior feature differences between the generated samples and the real vulnerability scenarios.
[0078] The output results of the discriminator network are fed back to the federated learning nodes through the gradient backpropagation mechanism to optimize the feature extraction rules of the distributed nodes. Specifically, the behavior difference features between the generated samples and the real samples identified in the discrimination results will be used as reinforcement signals to guide the federated learning nodes to adjust the attention dimensions of the local feature extractors. For example, when the discriminator detects a significant difference in the memory out-of-bounds behavior features between the generated samples and the real samples, the federated learning nodes will increase the parsing weight of the memory operation mode in the local code structure feature extractor to enhance the representation ability of the subsequent encrypted aggregated features for memory security vulnerabilities.
[0079] In the above operations, the generative adversarial network and federated learning form a two-way optimization mechanism: the global feature statistics provided by federated learning guide the generator to capture cross-node vulnerability patterns, and the generated samples expand the diversity of training data; the comparison results of the discriminator are used to reverse-optimize the feature extraction rules of the federated learning nodes, improving the vulnerability representation accuracy of the subsequent global features. This collaborative mechanism realizes the spiral optimization of data generation and feature extraction through iterative training, effectively solving the problem of insufficient model generalization ability in small-sample scenarios while protecting data security, and enhancing the detection robustness of polymorphic vulnerabilities in complex supply chain environments.
[0080] Specifically, for the intelligent location method of software supply chain vulnerabilities with multi-modal feature fusion described in the present invention, the multi-head cross-attention mechanism and the graph attention network cooperate to perform the following operations. The component risk probability calculated by the graph attention network is used as the weight constraint condition of the multi-head cross-attention mechanism to suppress the correlation feature interaction between code structure nodes and low-risk dependency paths. The cross-modal consistency feature representation generated by the multi-head cross-attention mechanism corrects the edge weight parameters of the supply chain dependency graph and optimizes the quantization accuracy of the vulnerability propagation probability.
[0081] The operation process of the cooperation between the multi-head cross-attention mechanism and the graph attention network in the present invention is realized through the following technical solutions: when the component risk probability calculated by the graph attention network is used as the weight constraint condition of the multi-head cross-attention mechanism, the component version vulnerability score and the dependency chain propagation path feature are extracted from the dependency graph nodes to generate a risk probability vector. This vector is converted into an attention mask matrix through normalization processing and acts on the weight calculation process of the multi-head cross-attention mechanism to suppress the feature interaction intensity between code structure nodes and low-risk dependency paths. For example, for function nodes that call components with low-risk versions, their priority in attention weight allocation is reduced to reduce the interference of irrelevant features on cross-modal alignment.
[0082] When the cross-modal consistency feature representation generated by the multi-head cross-attention mechanism corrects the edge weight parameters of the supply chain dependency graph, the feature projection technology is used to map the fused cross-modal vector to the edge feature space of the dependency graph. According to the projection results, the edge weight coefficients of the component call relationships in the dependency graph are adjusted. For example, the weight of the dependency path with a historical vulnerability propagation record is enhanced, and the edge weight of the component call without a discovered associated vulnerability is reduced. After the edge weights are corrected, the reachability probability of the vulnerability in the multi-layer dependency chain is recalculated to optimize the adaptability of the vulnerability propagation path quantization model to multi-modal features. For example, it can more accurately identify the indirect dependency risk path hidden through text description.
[0083] In the above operations, the interaction between the component risk probability and the cross-modal features forms a two-way optimization mechanism: the risk probability provided by the graph attention network constrains the feature alignment direction of the multi-head attention mechanism, reducing the noise interference of low-risk paths; the cross-modal features feed back to the topological structure analysis of the dependency graph through edge weight correction, enhancing the context awareness ability of vulnerability propagation probability calculation. The multi-modal data of code structure, dependency relationship, and dynamic behavior are hierarchically fused through this cooperation mechanism, and the static topological analysis of the dependency graph and the dynamic semantic perception of cross-modal features are mutually enhanced, solving the detection blind spots of traditional methods for hidden dependency chains and cross-language interface vulnerabilities, and improving the comprehensiveness and accuracy of vulnerability propagation path analysis in complex supply chain scenarios.
[0084] Specifically, the intelligent software supply chain vulnerability localization method with multi-modal feature fusion according to the present invention further includes: Taking the vulnerability propagation probability of the identified high-risk dependency path as the prior knowledge of the state space of the differentiable reinforcement learning framework to dynamically narrow the scanning range of the target code; Based on the vulnerability localization result output by the differentiable reinforcement learning framework, the node risk features of the graph attention network are updated in reverse to form a closed-loop optimization of the dependency graph risk score and the localization strategy.
[0085] The operation process of the closed-loop optimization of the dependency graph and reinforcement learning according to the present invention is realized through the following technical solutions: when taking the vulnerability propagation probability of the identified high-risk dependency path as the prior knowledge of the state space of the differentiable reinforcement learning framework, feature extraction is performed on the component call chain marked as high-risk in the dependency graph to generate a path reachability probability vector. This vector is concatenated with the cross-modal association features output by the multi-head cross-attention mechanism to construct the multi-dimensional input features of the reinforcement learning state space. The scanning range of the target code is dynamically adjusted through a probability threshold screening mechanism. For example, fine-grained scanning is enabled for the code modules associated with the dependency paths whose reachability probability exceeds the preset threshold, and sampling detection is adopted for the regions associated with low-probability paths to optimize the allocation efficiency of computing resources.
[0086] When the node risk features of the graph attention network are updated in reverse based on the vulnerability localization result output by the differentiable reinforcement learning framework, the gradient backpropagation mechanism is used to feedback the false positive and false negative signals in the localization result to the dependency graph nodes. Specifically, for the code segments confirmed as vulnerabilities in the localization result, according to the actual vulnerability propagation situation of their associated dependency paths, the risk scores of the corresponding path nodes are enhanced; for false positive cases, the risk probability weights of the relevant nodes are adjusted downward according to the deviation degree between the code context semantics and the dependency paths. The updated node risk features are re-input into the graph attention network to correct the edge weights of the dependency graph, forming an iterative optimization closed-loop of the risk score and the localization strategy.
[0087] In the above operations, the interaction between the vulnerability propagation probability and the localization result forms a two-way optimization mechanism: the dependency graph provides prior knowledge to guide the generation of the localization strategy of the reinforcement learning, and the localization result corrects the risk assessment model of the dependency graph through gradient backpropagation. The dynamic adjustment of the scanning range reduces the ineffective computational overhead in low-risk areas, and the gradient feedback mechanism continuously optimizes the accuracy of the graph node features. This closed-loop design realizes the co-evolution of the dependency risk quantification and the vulnerability localization strategy, adapts to the vulnerability detection requirements in the scenario of continuous iteration of the software supply chain, and improves the detection rate and detection efficiency of long-tail vulnerabilities in complex dependency ecosystems.
[0088] Specifically, the intelligent software supply chain vulnerability localization method with multi-modal feature fusion according to the present invention further includes: The prototype network is used to cluster the vulnerability features of multi-modal training samples, and a prototype vector of vulnerability patterns across programming language projects is constructed. The localization strategy parameters of the differentiable reinforcement learning framework are transferred to the lightweight interpretation model through the attention distillation technology, and heatmap interpretation information corresponding to the vulnerable code segment and the associated dependency chain is generated.
[0089] The operation process of the cooperation between the prototype network and the attention distillation in the present invention is realized through the following technical solutions: when using the prototype network to cluster the vulnerability features of multi-modal training samples, the control flow anomaly pattern in the code structure diagram, the temporal anomaly features of the dynamic behavior sequence, and the semantic risk keywords in the text description are extracted, and the cosine similarity of the vulnerability features between different programming language projects is calculated through metric learning. The vulnerability features with similarity exceeding the preset threshold are clustered into a unified prototype vector. For example, the prototype vector of memory safety vulnerabilities aggregates the pointer misuse features of C / C++ projects and the lifetime management defect features of Rust projects to form a general representation of cross-language vulnerability patterns.
[0090] After constructing the prototype vector of the vulnerability pattern across programming language projects, the localization strategy parameters of the differentiable reinforcement learning framework are transferred to the lightweight interpretation model through the attention distillation technology. Specifically, the cross-modal attention weight distribution and the localization threshold decision rule are extracted from the policy network of the reinforcement learning framework and compressed into the classifier parameters of the lightweight model through the knowledge distillation algorithm. When the lightweight interpretation model performs vulnerability pattern matching on the target code segment based on the transferred parameters and generates heatmap interpretation information, the code lines and the associated dependency chain nodes with the highest matching degree with the prototype vector are highlighted through the gradient-weighted class activation mapping technology. For example, the transmission path of unvalidated input parameters and the call positions of high-risk dependent components are identified.
[0091] In the above operation, the prototype network and the attention distillation form a cooperative mechanism for cross-language vulnerability detection and interpretation: the prototype vector provides a benchmark for vulnerability patterns across projects, guiding the lightweight model to quickly adapt to the detection scenarios of new programming languages; the distillation technology retains the complex decision logic of reinforcement learning, and visualizes the key feature association paths through heatmaps. This mechanism realizes the interpretability verification of the vulnerability localization results. Security experts can optimize the clustering accuracy of the prototype vector through the heatmap annotation feedback, forming an interactive optimization closed-loop between the detection model and the expert experience, and improving the transparency and credibility of cross-ecosystem supply chain vulnerability governance.
[0092] The intelligent positioning method for software supply chain vulnerabilities with multimodal feature fusion provided by the present invention is implemented through the following steps in specific implementation: First, multimodal data collection is performed on the code structure data, dependent component metadata, and runtime behavior logs of the target software system, and federated learning framework is used to encrypt and aggregate the data of distributed nodes. Specifically, homomorphic encryption transmission is performed on the code syntax tree features, dependent version hash values, and runtime memory access sequences through a secure multi-party computing protocol, and weighted average operations are performed in ciphertext state on the central server to generate global feature statistics. During this process, the reinforcement learning agent analyzes the code change frequency based on the commit records of the version control system, quantifies the risk levels of dependent components in combination with CVSS scores, dynamically adjusts the node sampling weights, and preferentially collects the code features of frequently modified modules and the metadata of high-risk dependent components to improve the data collection efficiency and risk focusing ability.
[0093] When the encrypted feature set is input into the conditional generative adversarial network for data augmentation, the generator network generates code variants that conform to the abstract syntax tree rules under the constraint of the code structure parser, such as constructing functions with unvalidated input parameters or loop structures with memory out-of-bounds risks. The discriminator network compares the dynamic behavior sequences of the generated code variants with those of real vulnerability samples, extracts temporal feature differences such as memory operation frequencies and system call path similarities through a temporal convolutional layer, constructs an adversarial loss function to iteratively optimize the generation strategy, and outputs multimodal training samples with consistent semantics. The generated samples and real data jointly participate in subsequent training to enhance the robustness of the model against code obfuscation and adversarial attacks.
[0094] The multimodal training samples are input into the multi-head cross-attention mechanism for cross-modal feature alignment. The function call relationships in the code structure diagram are transformed into vector representations through graph embedding technology, combined with the pre-trained language model to encode text semantic features and the temporal convolutional network to extract dynamic event sequence features. The hierarchical attention mechanism calculates the interaction weights of code nodes, dependent components, and dynamic behaviors in multiple subspaces, suppresses the noise interference of low-risk paths, and generates a fused feature vector. This vector is input into the graph attention network to construct a supply chain dependency graph, dynamically integrate component version information and vulnerability disclosure records, and update the node risk features. The reachability probability of vulnerabilities in the multi-layer dependency chain is quantified through the attention weight propagation algorithm, such as identifying the propagation path of high-risk third-party libraries affecting the main program entry through multi-level call chains.
[0095] When using the differentiable reinforcement learning framework for vulnerability location, a state space is constructed based on reachability probability and cross-modal attention weights, integrating code context semantics, historical false positive statistics, and dependency risk scores. A reward function is set to balance the location accuracy and false positive suppression effect, and the scanning threshold is dynamically adjusted through the policy gradient algorithm. Fine-grained detection is preferentially enabled for code modules associated with high-risk dependency paths. The location results are fed back to the graph attention network through the gradient backpropagation mechanism to correct the node risk scores and edge weight parameters, forming a closed loop of "risk quantification - location - graph update". For example, when false positive cases show that the risk score of a certain dependency path is inflated, its weight is automatically lowered and the graph topology is updated to improve the subsequent detection accuracy.
[0096] In addition, a prototype network is used to perform cross-language clustering on the vulnerability features of multi-modal training samples. For example, the common features of C / C++ pointer misuse and Rust lifetime defects are aggregated to construct a prototype vector of vulnerability patterns across programming languages. The location policy parameters of reinforcement learning are migrated to a lightweight interpretation model through attention distillation technology to generate heatmap interpretation information, highlighting high-risk code segments and associated dependency chain nodes to assist security experts in verification and model optimization. The above steps achieve accurate location and cross-ecosystem generalization of complex supply chain vulnerabilities through a collaborative mechanism of data collection, feature fusion, dependency analysis, and dynamic decision-making, while protecting data security.
[0097] Example 1: In this example, the federated learning framework uses a secure multi-party computation protocol to encrypt and aggregate the code structure data, dependency component metadata, and runtime behavior logs of distributed nodes. The specific parameter settings are as follows: Encryption algorithm: The Paillier homomorphic encryption algorithm is used to process the code syntax tree feature vectors, and the key length is 2048 bits; Data aggregation: The weighted average operation in the ciphertext state is performed on the central aggregation server, and the weight coefficient is dynamically adjusted to 0.3 - 0.7 according to the node data volume; Dynamic sampling strategy: The node sampling weights are optimized based on the Q-learning algorithm. The code change frequency threshold is set to the number of daily commits ≥ 5 times, and the risk level of dependency components uses the CVSS score ≥ 7.0 as the high-risk threshold; Feature dimension: The dimension of the code structure feature vector is 512, the length of the dependency metadata hash value is 256 bits, and the runtime behavior log is encoded as a 128-dimensional time series vector.
[0098] Example 2: In this example, the parameter configuration of the generator network of the conditional generative adversarial network is as follows: Code Structure Parser: Constructed based on abstract syntax tree rules, supporting syntax verification for languages such as Java / Python / C++ etc., with the error tolerance rate for function closure check ≤ 0.1%; Generator Structure: Adopts a 5-layer Transformer encoder with a hidden layer dimension of 768, and the syntax compliance rate of the generated code variants ≥ 98%; Discriminator Comparison Parameters: For dynamic behavior sequence comparison, a temporal convolutional layer is used (kernel size = 5, stride = 2), the cosine similarity threshold is set to 0.85, and the number of adversarial training iterations is 1000 rounds; Output Samples: The vulnerability semantic matching degree of the generated code variants ≥ 90%, and the number of samples generated in a single batch is 512.
[0099] Example 3: In this example, the specific parameters of the multi-head cross-attention mechanism include: Number of Attention Heads: Set to 8 heads, with the dimension of each head of attention being 64; Cross-modal Association Weight Calculation: The vector dot product scaling factor between code structure nodes and dependency graph nodes is √64, and the association weight threshold ≥ 0.6; Hierarchical Attention Interaction: The text entity encoding uses the BERT pre-trained model (Base version), the time window size of the dynamic event sequence is 10 steps, and the dimension after feature concatenation is 1024; Fusion Feature Output: The dimension of the cross-modal consistency feature vector is 256, and the feature redundancy is reduced to less than 15%.
[0100] Example 4: In this example, the parameter configuration of the graph attention network is as follows: Node Risk Feature Update: For nodes where the component version release time exceeds 2 years and there is no maintenance record, the node risk weight increases by 0.3, and for nodes with a CVSS score ≥ 9.0, the risk coefficient is superimposed by 1.5; Reachability Probability Calculation: The number of iterations of the attention propagation algorithm is 3 rounds, the initial value of the edge weight is the normalized value of the component call frequency (in the range of 0 - 1), and the path reachability probability error rate ≤ 5%; Graph Optimization: When projecting cross-modal features into the edge weight space, a fully connected layer is used (dimension 128 → 64), and the edge weight correction step size is 0.01.
[0101] Example 5: In this example, the parameters of the differentiable reinforcement learning framework include: State Space Construction: The dimension of the code context semantic vector is 256, the time span of historical false alarm statistics is 30 days, and the dynamic update period of the dependency risk score is 1 hour; Reward function design: The recall weight coefficient is 0.7, the false alarm penalty term coefficient is 0.3, and the positioning accuracy deviation tolerance ≤ 2%; Policy gradient optimization: The Monte Carlo sampling times is 500 times, the learning rate is set to 0.001, and the convergence iteration times of the positioning threshold decision parameter is 200 rounds; Closed-loop feedback: The gradient backpropagation step size of the node risk feature update is 0.005, and the response delay of the dependency graph edge weight correction ≤ 1 second.
[0102] Example 6: In this example, the parameter configurations of the prototype network and the attention distillation technology are as follows: Vulnerability feature clustering: Using cosine similarity measurement, the clustering threshold is set to 0.8, and the cross-language vulnerability prototype vector dimension is 128; Knowledge distillation: The compression rate of the reinforcement learning policy network parameters is 60%, the number of parameters of the lightweight interpretation model ≤ 1M, and the response time of the heatmap generation ≤ 50ms; Interpretability verification: The focus area radius of the gradient-weighted class activation mapping is set to 3 lines of code, and the ratio of the heatmap resolution to the number of code lines is 1:1.
[0103] Explanation of the technical features of the present invention: Federated learning framework: Definition: A distributed machine learning paradigm that allows multiple participants to collaboratively train a model without the data leaving the local nodes.
[0104] Application in the present invention: Encrypted aggregation: Encrypt and transmit and aggregate the code structure data (such as abstract syntax tree features), dependency component metadata (version hash values), and runtime behavior logs (memory access sequences) of distributed nodes through a secure multi-party computing protocol (such as Paillier homomorphic encryption) to generate global feature statistics (such as 512-dimensional vectors).
[0105] Dynamic sampling optimization: Analyze the code change frequency (such as the number of daily commits ≥ 5 times) and the risk level of dependency components (CVSS score ≥ 7.0) based on a reinforcement learning agent (Q-learning algorithm), dynamically adjust the node sampling weights (weight coefficient 0.3 - 0.7), and preferentially collect data of high-frequency modified modules and high-risk dependency components. Technical effect: Solve the problem of multi-source data security protection and improve the pertinence and efficiency of data collection.
[0106] Conditional generative adversarial network (cGAN): Definition: A variant of the generative adversarial network that constrains the semantic attributes of the generated samples through conditional input. Application in the present invention: Generator Network: It takes the global feature statistics of federated learning as input, and generates syntactically compliant code variants (such as functions with unvalidated parameters) through a code structure parser (based on abstract syntax tree rules), with a compliance rate ≥ 98%.
[0107] Discriminator Network: It compares the dynamic behavior sequences of the generated code variants and real vulnerability samples (such as the system call chain of memory out-of-bounds operations), uses a temporal convolutional layer (kernel size = 5) to extract temporal feature differences, and optimizes the generation strategy through an adversarial loss function. Technical effect: It expands the small-sample training data and enhances the model's robustness against code obfuscation attacks.
[0108] Multi-Head Cross-Attention Mechanism: Definition: An attention mechanism for parallel computing of multi-modal feature correlation weights, which separates the interactions in different semantic spaces through a multi-head design. Application in the present invention: Cross-Modal Alignment: It parallelly computes the correlation weights (threshold ≥ 0.6) of function call nodes (graph embedding vectors) in the code structure graph, component nodes (risk probability vectors) in the dependency graph, and dynamic event sequences (temporal convolutional features), and generates a cross-modal consistent feature representation (256-dimensional vector).
[0109] Hierarchical Interaction: The text entity description is encoded by the BERT model (768-dimensional), the dynamic event sequence extracts temporal features through a sliding window (10 steps), and the multi-head attention (8 heads) hierarchically aggregates the features. Technical effect: It eliminates the semantic gap between code, text, and behavior data and reduces the interference of redundant features.
[0110] Graph Attention Network (GAT): Definition: A graph neural network based on the attention mechanism, used to quantify the association strength between nodes. Application in the present invention: Node Risk Update: According to the component version information (such as no maintenance for more than 2 years since the release time) and vulnerability disclosure records (CVE number, CVSS score ≥ 9.0), it dynamically updates the risk features of the dependency graph nodes (risk weight + 0.3).
[0111] Propagation Probability Calculation: It iteratively calculates the reachability probability of vulnerabilities in the dependency chain through the attention propagation algorithm (error rate ≤ 5%), and the edge weights are normalized according to the call frequency (in the range of 0 - 1). Technical effect: It accurately quantifies the vulnerability cross-component propagation path and supports the risk analysis of the dependency chain.
[0112] Differentiable Reinforcement Learning Framework (DRL): Definition: A decision-making framework that combines reinforcement learning and gradient optimization, supporting end-to-end policy tuning. Application in the present invention: State Space Construction: Integrate the code context semantic vector (256 - dimensional), historical false positive statistics (30 - day window), and dependency risk scores to generate a multi - dimensional state input.
[0113] Dynamic Policy Optimization: The reward function balances recall (weight 0.7) and false positive suppression (weight 0.3), and adjusts the positioning threshold (converging and iterating 200 rounds) through the policy gradient algorithm (learning rate 0.001). Technical Effect: Adaptively adjust the scanning granularity to balance detection efficiency and accuracy.
[0114] Prototype Network and Attention Distillation: Definition: Prototype Network: Construct prototype vectors (128 - dimensional) of cross - language vulnerability patterns through clustering to support few - shot migration.
[0115] Attention Distillation: Transfer the decision - making logic of complex models to lightweight models to generate interpretable heatmaps. Application in the present invention: Cross - language Clustering: Measure the cosine similarity (threshold ≥ 0.8) between C / C++ pointer misuse and Rust lifetime defects through metric learning to generate unified prototype vectors.
[0116] Heatmap Generation: Highlight high - risk code segments and dependency chain nodes (response time ≤ 50ms) through gradient - weighted class activation mapping (focus radius 3 lines of code).
[0117] Technical Effect: Improve cross - language generalization ability and enhance the interpretability of detection results.
[0118] Federated Learning Framework: Used to securely aggregate multi - modal data among distributed nodes. This framework encrypts and aggregates code structure data, dependency component metadata, and runtime behavior logs through a secure multi - party computation protocol to generate global feature statistics. The dynamic sampling strategy optimizes the node data weights based on a reinforcement learning agent, preferentially collecting data on frequently modified modules and high - risk dependency components to improve data coverage efficiency. The global feature statistics provide cross - node vulnerability pattern features for subsequent models while ensuring data security.
[0119] Conditional Generative Adversarial Network (cGAN): Includes a generator and a discriminator network. The generator receives the global feature statistics of federated learning and generates syntactically compliant code variants under the constraint of a code structure parser, such as constructing interfaces with unvalidated inputs or memory - out - of - bounds structures. The discriminator compares the dynamic behavior sequences (such as system call chains) of the generated code variants and real vulnerability samples, and optimizes the semantic authenticity of the generated samples through adversarial training. The generated code variants expand the multi - modal training samples, alleviate the data scarcity problem in few - shot scenarios, and enhance the model's robustness to code obfuscation.
[0120] Multi-Head Cross-Attention Mechanism: Cross-modal features for fusing code structure, text semantics, and dynamic behavior. Code structure nodes are mapped to vectors through graph embedding, text entities are encoded using pre-trained language models, and dynamic event sequences extract temporal features through temporal convolution. Parallel attention subspaces calculate interaction weights between different modalities, such as the association strength between code function nodes and abnormal memory operations. Hierarchical attention aggregates multi-modal features to generate a unified fusion vector, eliminating the semantic gap and providing multi-dimensional input for dependency graph analysis.
[0121] Graph Attention Network (GAT): Construct a supply chain dependency graph and quantify the vulnerability propagation path. Node risk attributes are updated based on component version information and vulnerability disclosure records, and edge weights reflect component call frequency and interface complexity. The reachability probability of vulnerabilities in multi-layer dependency chains is calculated through the attention propagation algorithm, and the edge weight parameters are corrected by combining cross-modal fusion features to identify high-risk dependency paths (such as call chains of high-risk third-party libraries). The reachability probability is used as an input parameter for reinforcement learning to guide the vulnerability location priority.
[0122] Differentiable Reinforcement Learning Framework (DRL): Dynamically adjust the vulnerability location strategy based on the state space. The state space integrates code context semantics, historical false positive statistics, and dependency risk scores, and the reward function balances recall and false positive suppression. The policy gradient algorithm optimizes the scanning threshold parameter, preferentially scans code modules associated with high-risk dependency paths, and outputs vulnerable code segments and complete dependency chains. The location results are fed back to the graph attention network to form a closed loop of "risk quantification - location - graph update", improving the detection adaptability.
[0123] Prototype Network and Attention Distillation Technology: The prototype network clusters cross-language vulnerability features through metric learning to construct a prototype vector of common vulnerability patterns (such as common features of memory security vulnerabilities). Attention distillation transfers the reinforcement learning policy parameters to a lightweight model to generate a heatmap to explain vulnerable code segments and associated dependency chains, such as highlighting code lines with unvalidated input parameters. This technology enhances the generalization ability in cross-language scenarios and improves the interpretability of detection results through visualization to assist security experts in verification.
[0124] The above models form a collaborative closed loop through multi-modal data collection, feature fusion, dependency analysis, and dynamic decision-making. Federated learning ensures data security, generative adversarial networks enhance sample diversity, multi-head attention achieves cross-modal alignment, graph attention networks quantify propagation risks, reinforcement learning optimizes location strategies, and prototype networks and distillation technologies support cross-language interpretation. The technical connection and feedback mechanism of each model systematically solve the deficiencies of traditional methods in data security, small-sample generalization, dependency chain analysis, and interpretability, and achieve precise location and governance of complex supply chain vulnerabilities.
[0125] The present invention systematically solves the problems of the prior art through the following technical solutions: Aiming at the lack of vulnerability characterization dimensions caused by single-modal feature analysis and the difficulty of cross-modal fusion, the present invention uses a federated learning framework to encrypt and aggregate code structures, dependent component metadata, and runtime behavior logs to generate cross-node encrypted feature sets, realizing the protection and preliminary fusion of multi-source heterogeneous data. Through a conditional generative adversarial network, semantically realistic code variants are generated under the constraint of grammar rules, and the discriminator network is combined to compare the similarity of dynamic behavior sequences to expand multi-modal training samples. The multi-head cross-attention mechanism further extracts cross-modal correlation features of code structure nodes, text entity embedding vectors, and dynamic event sequences, eliminates semantic gaps through hierarchical attention interaction, and generates fused feature vectors. The present invention breaks through the limitations of single-modal analysis, reduces redundancy and noise interference through multi-dimensional feature alignment, and improves the comprehensiveness of vulnerability characterization.
[0126] Aiming at the scarcity of samples and the lack of cross-language generalization ability, the present invention introduces a prototype network to cluster the vulnerability features of multi-modal training samples and constructs prototype vectors of vulnerability patterns for cross-programming language projects. By measuring learning to quantify the similarity of vulnerability features of projects in different languages (such as the commonality between C / C++ pointer misuse and Rust lifetime defects), a cross-language vulnerability general representation is formed. Combining attention distillation technology, the positioning strategy parameters of the reinforcement learning framework are transferred to a lightweight interpretation model to generate heatmap interpretation information to guide the model to quickly adapt to new language scenarios. The collaborative mechanism of federated learning and generative adversarial networks uses global feature statistics to guide data augmentation, improving the model's robustness to noise and cross-ecosystem generalization ability under small sample conditions.
[0127] Aiming at the insufficient fusion of dependency graphs and multi-modal data and the limitations of vulnerability propagation path analysis, the present invention uses a graph attention network to construct a supply chain dependency graph, quantifies the vulnerability propagation probability between components, and inputs the fused feature vectors into the graph analysis module to identify high-risk dependency paths. The differentiable reinforcement learning framework constructs a state space based on the propagation probability and cross-modal attention weights, and dynamically adjusts the positioning threshold strategy. Through the gradient backpropagation mechanism, the positioning results are fed back to update the risk features of the dependency graph nodes, forming a closed loop of "risk quantification - positioning - graph update". The present invention breaks through the local association limitations of traditional dependency chains, combines multi-modal features to dynamically correct edge weight parameters, accurately tracks cross-level vulnerability diffusion paths, and improves the positioning accuracy and real-time response ability in complex supply chain environments.
Claims
1. The method for intelligently locating software supply chain vulnerabilities based on multimodal feature fusion is characterized by: include: Collect multimodal data of the target software system, encrypt and aggregate the code structure data, dependent component metadata, and runtime behavior logs of distributed nodes through the federated learning framework, and generate a cross-node encrypted feature set; The encrypted feature set is input into a conditional generative adversarial network, a code variant semantically consistent with the real vulnerability scenario is generated through the generator network, and the similarity of dynamic behavior sequences is compared through the discriminator network to output a multimodal training sample; Input the multimodal training samples into a multi-head cross attention mechanism, extract cross-modal correlation features between function call nodes, text entity embedding vectors and dynamic event sequence encodings in the code structure diagram, and generate a fused feature vector; Based on the dependency component metadata, a supply chain dependency graph is constructed through a graph attention network to quantify the probability of vulnerability propagation between components, and the fused feature vector is input into the graph attention network to identify high-risk dependency paths associated with code structure nodes in the dependency graph; A differentiable reinforcement learning framework is adopted to perform serialized scanning on the target code based on the vulnerability propagation probability of the high-risk dependency path and the attention weight of the cross-modal association feature, dynamically adjust the vulnerability location threshold strategy, and output the vulnerability code segment and the associated dependency chain.
2. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: The federated learning framework includes: encrypting and aggregating the code structure data, dependency component metadata, and runtime behavior logs through a secure multi-party computing protocol to generate global feature statistics; The frequency of changes in the code structure data and the risk level of the metadata of dependent components are analyzed based on the reinforcement learning agent, and the data sampling weight of each distributed node is dynamically adjusted to prioritize the collection of code features of high-frequency change modules and version metadata of high-risk dependent components.
3. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: The conditional generative adversarial network comprises: inputting global feature statistics generated by the federated learning framework into a generator network, and constraining the generated code variants to conform to abstract syntax tree rules through a code structure parser; The generated code variants and dynamic behavior sequences are input into the discriminator network, and the generator network is iteratively optimized to generate semantically consistent multimodal training samples.
4. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: The multi-head cross attention mechanism includes: based on the fused feature vector, parallel calculation of cross-modal association weights of function call nodes in the code structure graph and component nodes in the supply chain dependency graph; The embedding vectors of the code structure nodes, the text entity description encoding and the dynamic event sequence features are subjected to hierarchical attention interaction to generate a cross-modal consistent feature representation.
5. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: The graph attention network includes: updating the node risk characteristics of the supply chain dependency graph based on the version information and vulnerability disclosure records of the dependency component metadata; calculating the accessibility probability of the vulnerability in the multi-layer dependency chain through the attention weight propagation algorithm, and using the accessibility probability as the positioning strategy input parameter of the differentiable reinforcement learning framework.
6. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: The differentiable reinforcement learning framework includes: constructing a state space including a code context semantic vector, historical false positive statistics, and dependency risk scores based on the vulnerability propagation probability of the high-risk dependency path and the attention weight of the cross-modal association features; A reward function is set to evaluate the balance between vulnerability localization accuracy and false alarm suppression effect, and the localization threshold decision parameters of the state space are optimized through a policy gradient algorithm.
7. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 2 is characterized in that: Also includes: Using the generated global feature statistics as a generator input of a conditional generative adversarial network, constraining the generated code variants to be semantically consistent with the global feature statistics; The discriminator network compares the dynamic behavior sequences of the generated adversarial samples with those of the real vulnerability samples, and transmits the discrimination results back to the federated learning nodes to optimize the code structure feature extraction rules of the distributed nodes.
8. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 4 is characterized in that: Also includes: The component risk probability calculated by the graph attention network is used as the weight constraint of the multi-head cross attention mechanism to suppress the interaction of the associated features between the code structure nodes and the low-risk dependency paths; The cross-modal consistency feature representation generated by the multi-head cross-attention mechanism is used to correct the edge weight parameters of the supply chain dependency graph and optimize the quantification accuracy of the vulnerability propagation probability.
9. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 5 is characterized in that: Also includes: The vulnerability propagation probability of the identified high-risk dependency path is used as the state space prior knowledge of the differentiable reinforcement learning framework to dynamically narrow the scanning scope of the target code; The vulnerability localization results output by the differentiable reinforcement learning framework are used to reversely update the node risk features of the graph attention network.
10. The method for intelligently locating software supply chain vulnerabilities by multimodal feature fusion according to claim 1 is characterized in that: Also includes: The prototype network is used to cluster the vulnerability features of multimodal training samples and construct vulnerability pattern prototype vectors across programming language projects; The localization strategy parameters of the differentiable reinforcement learning framework are migrated to a lightweight interpretation model through the attention distillation technology to generate heat map interpretation information corresponding to the vulnerable code segment and the associated dependency chain.
Citation Information
Patent Citations
Smart contract vulnerability detection method based on directed graph attention network
CN117725592A
Multimodal source code vulnerability static detection method based on Transform and GAT
CN119203149A
Multi-source software supply chain intelligent analysis method and system
CN119720225A
Systems and methodologies for auto labeling vulnerabilities
US20250021657A1
Cited By
Software code security protection method and system based on domestic platform
CN120234801A
Software code security protection method and system based on domestic platform
CN120234801B
Software defect collaborative detection method and system based on multi-agent dynamic adaptation
CN120560989A
Small sample event detection method based on dynamic semantic element body driving
CN120561694A
Intelligent comparison and analysis method and system for similarity of examination answer codes
CN120631735A