A Multi-Source Software Supply Chain Intelligent Analysis Method and System

Through the intelligent analysis method of multi-source software supply chain, a weighted dependency relationship chain is built, a Bayesian network and anomaly detection algorithm are used to evaluate risks, and a risk map is generated through multi-dimensional visualization, which solves the problem of insufficient comprehensive and flexible risk assessment in the existing technology, and achieves accurate risk management and mitigation of the software supply chain.

CN119720225BActive Publication Date: 2025-06-20HUAQING WEIYANG (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510222232.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-20
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing software supply chain risk assessment solutions have problems with insufficient ability to identify new attacks and unknown threats, neglect indirect dependencies, lack of flexibility and adaptability, and difficulty in intuitively displaying complex dependencies and risk patterns.

Method used

The intelligent analysis method of multi-source software supply chain is adopted to analyze the identification information of software components, extract attribute characteristics, build a weighted dependency chain list, calculate risk scores using Bayesian network model and anomaly detection algorithm, identify potential threat points, and generate a comprehensive risk map through multi-dimensional visualization algorithm to provide risk mitigation strategies.

Benefits of technology

It realizes a comprehensive risk assessment of the software supply chain, accurately identify potential threat points, improves security protection capabilities, provides intuitive risk distribution display and specific risk mitigation strategies, and enhances the security and integrity of the software development, distribution and deployment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119720225B_ABST
    Figure CN119720225B_ABST
Patent Text Reader

Abstract

This application provides a multi-source software supply chain intelligent analysis method and system. Among them, it parses the software component identification information submitted by the user, extracts attribute features, locates the node paths between components, and generates a weighted dependency chain list containing direct and indirect dependency chains; uses a Bayesian network model combined with a predefined risk index system to calculate the risk score of each dependency chain, and identifies potential supply chain threat points through an anomaly detection algorithm to generate a risk report; this report covers the overall risk score and threat points of key links or components; based on this report, uses a multi-dimensional visualization algorithm to draw a comprehensive risk map to display groups of components with similar risk patterns; fuses genetic algorithms and simulated annealing algorithms to explore high-risk areas, proposes adjustment suggestions, and generates risk mitigation strategies to reduce the overall risk score. The technical solution provided by the embodiments of this application improves the security and integrity of the multi-source software supply chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of software engineering, and in particular, to a multi-source software supply chain intelligent analysis method and system. Background Art

[0002] With the widespread use of open-source components and third-party libraries, the complexity of software systems has increased sharply, and the dependency chain has become intricate. Developers and enterprises need an intelligent analysis method that can parse the software component identification information submitted by users and extract attribute features therefrom to accurately locate the node paths between software components and generate a weighted dependency chain list including direct and indirect dependency chains.

[0003] Currently, software supply chain risk management mainly relies on static code analysis, vulnerability scanning, and dependency management tools. These tools detect known vulnerabilities and insecure dependencies by scanning source code and dependency libraries. Some advanced tools also incorporate machine learning algorithms to identify abnormal behaviors and potential threats.

[0004] Existing software supply chain risk assessment solutions have several significant drawbacks. First, traditional static analysis and vulnerability scanning tools can only identify known vulnerabilities and are powerless against new types of attacks and unknown threats. Second, these tools often ignore the indirect dependencies in the dependency chain, resulting in an incomplete and inaccurate risk assessment. Third, most existing solutions calculate risk scores based on predefined rules, lacking flexibility and adaptability, and unable to dynamically adjust weights and assessment criteria according to actual situations. Finally, existing visualization tools and risk mitigation measures are relatively simple, making it difficult to intuitively display complex dependency relationships and risk patterns, and even less able to provide systematic optimization strategies. Summary of the Invention

[0005] The embodiments of the present application provide a multi-source software supply chain intelligent analysis method and system to solve the problems of insufficient security and low integrity of multi-source software supply chains in the prior art.

[0006] In a first aspect, the embodiments of the present application provide a multi-source software supply chain intelligent analysis method, including:

[0007] Parsing the identification information of software components submitted by users, and extracting the attribute features of the software components from the identification information;

[0008] According to the attribute features of the software components, locating the node paths associated between the software components to generate the dependency chains in the node paths, assigning weight values to the dependency chains, and generating a weighted dependency chain list, where the dependency chains include direct dependency chains and indirect dependency chains;

[0009] Using a Bayesian network model combined with a predefined risk index system, calculate the risk scores of the dependency chains in the weighted dependency chain list, and use an anomaly detection algorithm to identify potential supply chain threat points to generate a risk report including the overall risk score and the supply chain threat points, where the supply chain threat points are the key links or components that can be exploited by attackers to undermine software integrity and security during software development, distribution, and deployment;

[0010] According to the risk report, use a multi-dimensional visualization algorithm to draw a comprehensive risk map to obtain a risk view, and fuse a genetic algorithm and a simulated annealing algorithm to explore adjustment suggestions for high-risk areas of the risk view to generate a risk mitigation strategy; where the comprehensive risk map groups and displays components with similar risk patterns through a hierarchical clustering algorithm, and the risk pattern is the result obtained by judging the similarity of the dependency chains, the supply chain threat points, and the risk scores, and the risk mitigation strategy is used to reduce the risk scores of the software components.

[0011] Optionally, the step of using a Bayesian network model combined with a predefined risk index system to calculate the risk scores of the dependency chains in the weighted dependency chain list includes:

[0012] According to the Bayesian network model, use the Node2Vec or GraphSAGE graph embedding algorithm to vectorize the node paths in the weighted dependency chain list, map the component nodes, edge attributes, and topological structures in the dependency chain to a low-dimensional vector space to obtain dependency vectors;

[0013] Based on the dependency vectors, normalize the vulnerability quantity, update frequency, and community activity indicators through a dynamic weight allocation strategy, where the vulnerability quantity weight is dynamically adjusted according to the severity level of the CVE database, the update frequency weight is calculated by exponential decay in combination with the component version iteration cycle, and the community activity weight is linearly interpolated based on the code submission frequency and issue response time to form a multi-dimensional weighted initial risk score;

[0014] Using the dependency chains with the multi-dimensional weighted initial risk scores assigned, take the risk probability of the parent nodes in the Bayesian network as the feature input of the gradient boosting decision tree, extract the component dependency depth, cross-ecosystem call relationship, and license conflict type as additional features through feature engineering, and use a greedy algorithm to optimize the conditional probability convergence direction when splitting nodes of the decision tree to obtain an optimized conditional probability;

[0015] According to the optimized conditional probability, the Stacking ensemble learning method is used to perform meta-model fusion on the component vulnerability prediction results of the random forest classifier and the risk propagation coefficient prediction results of the LightGBM regressor, determine the weights of each base model through cross-validation, and output the risk assessment score of the dependency chain.

[0016] Optionally, the use of anomaly detection algorithms to identify potential supply chain threat points to generate a risk report including the overall risk score and the supply chain threat points includes:

[0017] Based on the overall risk score, the Isolation Forest algorithm is used to perform unsupervised clustering on the topological feature vectors of the dependency chain. Combining the characteristics of supply chain attack patterns recorded in the historical attack sample library, the dynamic percentile method based on a sliding window is used to set the anomaly threshold, and when the vulnerability density of the component dependency path exceeds 3 times the standard deviation of the same type of components in the same period, it is marked as a threat point;

[0018] The BERT pre-trained model is used to perform entity recognition on security bulletins. Combining dependency syntactic analysis to extract triple information of CVE numbers, affected version ranges, and repair solutions, and using knowledge graph alignment technology to associate and match the extracted threat information with the component metadata in the software bill of materials to generate threat point details including vulnerability propagation paths;

[0019] All the overall risk scores and the threat point details are summarized and sorted. The Prophet time series model is used to perform seasonal decomposition on the component vulnerability disclosure frequency and the number of dependency relationship changes. Based on the LSTM neural network, a risk trend prediction model is constructed. Using the historical risk score, component update interval, and community discussion heat as input features, a risk fluctuation prediction curve for the next three version cycles is generated;

[0020] A visualization tool is used to construct a three-dimensional topology map through the ECharts framework to present the component dependency network, use a Sankey diagram to visualize the risk propagation path, and integrate the D3.js force-directed graph to achieve dynamic focus display of high-risk nodes, generating an interactive risk report including a risk trend overlay layer with a time dimension and a threat point heat distribution layer.

[0021] Optionally, for the dependency chain with the allocated multi-dimensional weighted initial risk score, the risk probability of the parent node in the Bayesian network is used as the feature input of the gradient boosting decision tree. Additional features such as component dependency depth, cross-ecosystem call relationship, and license conflict type are extracted through feature engineering. The greedy algorithm is used to optimize the conditional probability convergence direction when splitting nodes of the decision tree to obtain the optimized conditional probability, including:

[0022] Collect and organize the low-dimensional vector representations of the dependency chains and the initial risk scores, extract the component dependency path depth, cross-ecosystem call frequency, and license conflict type as core features, combine the component version iteration interval, vulnerability repair response time, and developer contribution activity to construct an extended feature set, standardize the features to eliminate dimensional differences, and screen out the feature subset that has a significant impact on risk propagation through feature importance analysis to generate a multi-dimensional feature training dataset;

[0023] Based on the Bayesian network model, initialize the network parameters through the correlation analysis of the risk status of the parent node and the dependency strength of the child node, use the non-linear fitting ability of the decision tree model to assist in establishing the initial mapping relationship of the conditional probability between nodes, and use the parent node risk probability, component dependency depth, and cross-ecosystem call frequency as the input features of the decision tree to construct a joint training framework;

[0024] Use the multi-dimensional feature training dataset, adopt an alternating optimization strategy to synchronously update the conditional probability table of the Bayesian network and the splitting rules of the decision tree, calculate the deviation between the model prediction risk probability and the true vulnerability disclosure record through the loss function, combine the backpropagation algorithm to dynamically adjust the network weights and tree structure parameters, and select the optimal splitting node according to the greedy algorithm in each iteration to optimize the convergence direction of the conditional probability until the model converges;

[0025] According to the joint training results of the joint training framework, weighted correction is performed on the conditional probabilities of the child nodes on the high-probability propagation paths in the Bayesian network, and the probability distribution parameters are refined based on the component version compatibility data in the dependency chain and the statistical results of historical attack events. The optimized conditional probability values are fed back into the Bayesian network to update the dependency relationship strength between nodes to generate optimized conditional probabilities.

[0026] Optionally, based on the overall risk score, use the isolation forest algorithm to perform unsupervised clustering on the topological feature vectors of the dependency chains, combine the supply chain attack pattern features recorded in the historical attack sample library, and use the dynamic percentile method based on a sliding window to set the anomaly threshold. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of the same type of components in the same period, it is marked as a threat point, including:

[0027] According to the overall risk score, extract the topological feature vectors of the dependency chains, including component node degree centrality, dependency path length, and cross-ecosystem call density, combine the supply chain attack pattern features recorded in the historical attack sample library, construct a multi-dimensional risk feature space, normalize the features to eliminate dimensional differences, and reduce the dimension through principal component analysis to improve the calculation efficiency, and generate an anomaly detection model adapted to the current software ecosystem;

[0028] Use an unsupervised clustering algorithm to group the topological feature vectors, combine the high-risk sample annotations in the historical attack sample library, train an anomaly detection model, optimize the algorithm parameters through model evaluation metrics, and generate a detection model that can identify outliers in the multi-dimensional risk feature space;

[0029] Based on the historical data of the dependency chain, statistically analyze the risk score distribution interval of similar components in the normal operation state, set the initial anomaly determination boundary in combination with the high-risk sample set annotated by domain experts, adopt the dynamic percentile method based on a sliding window, and monitor the vulnerability disclosure frequency in the component repository update log in real time. Dynamically expand the anomaly feature dimension according to the newly emerging attack vector features, and generate a dynamic threshold that matches the current threat environment;

[0030] Use the anomaly detection model combined with the dynamic threshold to analyze the risk score mutation phenomenon during the component version iteration in the dependency chain, identify composite threat points that simultaneously meet the topological structure anomaly, excessive vulnerability density, and lag in community repair response. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of the same type of components in the same period, mark it as a supply chain threat point and generate a supply chain threat point list.

[0031] Optionally, according to the risk report, use a multi-dimensional visualization algorithm to draw a comprehensive risk map to obtain a risk view, including:

[0032] According to the risk report, integrate the dependency chain topology data, threat point spatial location information, and risk score time series, and construct a structured data set containing component coordinates, risk level labels, and threat propagation direction vectors;

[0033] Use a multi-dimensional visualization algorithm to map the dependency chain into a connection line in three-dimensional space, convert the risk score into a node size parameter, and mark the threat point position as a pulse signal marker to generate an interactive spatial distribution map;

[0034] Based on the visual representation, use density-aware clustering technology to automatically delineate the high-risk component aggregation area, generate risk pattern groups according to the development team, license type, and vulnerability history similarity of the components, and realize the hierarchical display of the threat propagation path;

[0035] Combined with the network topology structure, design a circular layout algorithm to optimize the node arrangement, use gradient color coding to represent the risk trend changes in different time windows, and integrate the focus + context visualization technology to realize the scalable display of large-scale dependency networks, and generate a comprehensive risk map containing spatio-temporal multi-dimensional features;

[0036] Overlay and display the vulnerability infection intensity between components through a heat map layer, use a dynamic streamline diagram to present the diffusion path of risks in the supply chain, and provide a risk node penetration query function to reveal the underlying threat evidence chain to form an actionable risk view.

[0037] Optionally, based on the attribute characteristics of the software components, locate the node paths associated between the software components to generate a dependency relationship chain in the node paths, assign weight values to the dependency relationship chain, and generate a weighted dependency relationship chain list, including:

[0038] Utilize the parsed software component identification information submitted by the user to extract structured attribute characteristics including version number, build tool type, and dependency declaration file format;

[0039] According to the attribute characteristics of the software components, establish a multi-level index model of component dependencies, identify explicit declared dependencies and implicit derived dependencies through a dependency resolution engine, and generate a dependency relationship network including version constraint conditions;

[0040] Based on the dependency relationship network, trace the call records of third-party libraries dynamically loaded during the build process, analyze the interface call paths covered by test cases, and verify the actually effective dependency relationships in combination with continuous integration logs to generate a dependency relationship chain record accurate to specific call links;

[0041] Classify the dependency relationship chain records, divide direct and indirect dependencies according to the dimensions of whether the dependency introduction method is necessary for the development environment and whether it is forcibly loaded during runtime, and assign exponentially increasing propagation weights to deeply nested cross-ecosystem dependencies to generate a preliminary weighted dependency relationship chain;

[0042] Based on the comprehensive risk report, establish a weight correction rule engine, dynamically adjust the weight coefficients according to security factors such as the reputation rating of component maintainers, the verification status of binary file signatures, and the integrity of dependency lock files, and verify the rationality of weight assignment through dependency relationship propagation simulation, and finally generate a weighted dependency relationship chain list optimized through multiple rounds of iteration.

[0043] In a second aspect, an embodiment of the present application provides a multi-source software supply chain intelligent analysis system, including:

[0044] A parsing module, configured to parse the identification information of the software components submitted by the user and extract the attribute characteristics of the software components from the identification information;

[0045] A positioning module, configured to locate the node paths associated between the software components according to the attribute characteristics of the software components to generate a dependency relationship chain in the node paths, assign weight values to the dependency relationship chain, and generate a weighted dependency relationship chain list, where the dependency relationship chain includes a direct dependency relationship chain and an indirect dependency relationship chain;

[0046] A calculation module, configured to calculate the risk scores of the dependency chains in the weighted dependency chain list by using a Bayesian network model in combination with a predefined risk index system, and identify potential supply chain threat points by using an anomaly detection algorithm, so as to generate a risk report including the overall risk score and the supply chain threat points, wherein the supply chain threat points are key links or components that can be exploited by attackers to undermine the integrity and security of software during the software development, distribution, and deployment processes;

[0047] A plotting module, configured to draw a comprehensive risk map by using a multi-dimensional visualization algorithm according to the risk report to obtain a risk view, and explore adjustment suggestions for high-risk areas of the risk view by integrating a genetic algorithm and a simulated annealing algorithm, so as to generate a risk mitigation strategy; wherein the comprehensive risk map groups and displays components with similar risk patterns through a hierarchical clustering algorithm, and the risk pattern is a result obtained by judging the similarity of the dependency chains, the supply chain threat points, and the risk scores, and the risk mitigation strategy is used to reduce the risk scores of the software components.

[0048] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-source software supply chain intelligent analysis method as described in the first aspect above.

[0049] In a fourth aspect, an embodiment of the present application provides a computer storage medium, storing a computer program, and when the computer program is executed by a computer, a multi-source software supply chain intelligent analysis method as described in the first aspect is implemented.

[0050] In the embodiments of the present application, the identification information of the software components submitted by the user is parsed, and the attribute characteristics of the software components are extracted from the identification information; according to the attribute characteristics of the software components, the node paths associated between the software components are located to generate a dependency chain in the node paths, a weight value is assigned to the dependency chain, and a weighted dependency chain list is generated, where the dependency chain includes a direct dependency chain and an indirect dependency chain; a Bayesian network model is used in combination with a pre-defined risk index system to calculate the risk scores of the dependency chains in the weighted dependency chain list, and an anomaly detection algorithm is used to identify potential supply chain threat points to generate a risk report including the overall risk score and the supply chain threat points, where the supply chain threat points are key links or components that are exploited by attackers to undermine the integrity and security of the software during the software development, distribution, and deployment processes; according to the risk report, a multi-dimensional visualization algorithm is used to draw a comprehensive risk map to obtain a risk view, and a genetic algorithm and a simulated annealing algorithm are integrated to explore adjustment suggestions for high-risk areas of the risk view to generate a risk mitigation strategy; where the comprehensive risk map is grouped and displayed by a hierarchical clustering algorithm to show components with similar risk patterns, and the risk pattern is the result obtained by judging the similarity of the dependency chain, the supply chain threat points, and the risk scores, and the risk mitigation strategy is used to reduce the risk score of the software components.

[0051] The technical solution of the present application has the following beneficial effects:

[0052] By parsing the attribute characteristics and dependency chains of software components, a comprehensive risk assessment of the software supply chain is realized to ensure no omission. Using a Bayesian network model and an anomaly detection algorithm, potential threat points in the supply chain are accurately identified to improve the security protection ability. A multi-dimensional visualization algorithm is used to draw a comprehensive risk map to intuitively display the risk distribution, facilitating decision-makers to quickly understand complex risk patterns. A genetic algorithm and a simulated annealing algorithm are integrated to explore the best adjustment plan for high-risk areas to generate specific and feasible risk mitigation strategies, effectively reducing the risk score. Through systematic risk assessment and mitigation measures, the security and integrity of the software development, distribution, and deployment processes are enhanced, the attack surface is reduced, and the software quality is guaranteed.

[0053] Furthermore, the weighted dependency chain list is converted into a low-dimensional vector representation using a graph embedding algorithm to obtain dependency vectors; based on the dependency vectors, an initial risk score is calculated using a risk metric system that includes metrics such as the number of vulnerabilities, update frequency, and community activity, combined with a multi-dimensional weight adjustment mechanism, and this score is assigned to the dependency chains; using the dependency chains with the initial risk scores assigned, the risk probability distribution of the child nodes is optimized under the condition of a given software component parent node through a Bayesian network model combined with a gradient boosting decision tree algorithm to obtain the optimized conditional probability; according to the optimized conditional probability, combined with an ensemble learning method, the overall risk score of the dependency chains is evaluated to obtain the final overall risk score.

[0054] Through the above method, the accuracy and reliability of software supply chain risk assessment are significantly improved, the complex structure of the dependency chains is effectively captured, a precise low-dimensional vector representation is provided, the flexibility and adaptability of risk assessment are ensured, the risk probability distribution of each node in the dependency chains is optimized, the accuracy of risk prediction is improved, and the comprehensiveness and credibility of the assessment results are ensured.

[0055] Furthermore, based on the overall risk score, combined with an adaptive threshold setting strategy, an anomaly detection algorithm is used to analyze the risk scores of the dependency chains, identify data points that deviate from the normal range, determine the supply chain threat points, use natural language processing technology to parse the security bulletins and technical documents related to these threat points, automatically extract directly relevant information, generate detailed threat point profiles, summarize all the overall risk scores and threat point details, and apply time series analysis to predict future risk trends, form a risk report, and create an interactive interface through a visualization tool to display the risk report, generating a risk report that includes the overall risk score and threat point profiles.

[0056] Through the above method, the precision and practicality of software supply chain risk management are significantly enhanced, the false alarm or missed alarm problems caused by fixed thresholds are avoided, the efficiency and accuracy of information acquisition are improved, it helps decision-makers make preparations in advance, makes complex risk reports intuitive and easy to understand, facilitates decision-makers to quickly grasp the situation and take corresponding measures, and this method overall improves the quality and operability of the risk report, providing strong support for the security management of the software supply chain.

[0057] In summary, through intelligent analysis of the dependency chains and risk scores of the software supply chain, this method identifies and visualizes potential threat points, provides optimized risk mitigation strategies, and significantly improves the security and integrity of software development, distribution, and deployment.

[0058] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1 The flowchart of a multi-source software supply chain intelligent analysis method provided by the present application is shown;

[0061] Figure 2 The structural schematic diagram of a multi-source software supply chain intelligent analysis system provided by the present application is shown;

[0062] Figure 3 The structural schematic diagram of a computing device provided by the present application is shown. Detailed implementation manners

[0063] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application.

[0064] In some processes described in the specification and claims of the present application and the above-mentioned drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0066] Figure 1 A flowchart of a multi-source software supply chain intelligent analysis method is provided for the embodiments of the present application. As Figure 1 shown, the method includes:

[0067] 101. Analyze the identification information of the software components submitted by the user, and extract the attribute features of the software components from the identification information;

[0068] Among them, the identification information includes metadata such as the name, version number, release date, developer information, etc. of the software component, and these information are used to uniquely identify and describe a software component. The attribute features are specific characteristics further extracted from the identification information, such as the function description of the component, the list of dependent libraries, the type of open source license, etc. The process of analyzing the identification information and extracting the attribute features is the basis for ensuring the accuracy of subsequent analysis and provides key data input for subsequent steps.

[0069] In this solution, first, the system receives the software component identification information submitted by the user, and reads and processes this information through a parsing tool (such as a JSON parser or an XML parser). Then, using natural language processing (NLP) technology and pattern matching algorithms, the attribute features related to security and dependency relationships are extracted. These features will serve as the basic data for subsequent risk assessment and dependency relationship analysis.

[0070] For example, in an open source software management platform, the user uploads a new Python library. The system automatically parses the setup.py file of the library to obtain its identification information, such as the package name requests, the version number 2.25.1, the author Kenneth Reitz, etc. Then, through NLP technology, analyze the README.md document to extract the function description, the list of dependent libraries (such as urllib3, chardet, etc.) and the type of open source license (such as MIT License). These information provide an important basis for the generation of subsequent dependency chains.

[0071] 102. According to the attribute features of the software components, locate the node paths associated between the software components to generate the dependency chains in the node paths, assign weight values to the dependency chains, and generate a weighted dependency chain list, where the dependency chains include direct dependency chains and indirect dependency chains;

[0072] The node path refers to the dependency relationship link between software components, including direct dependencies (such as A depends on B) and indirect dependencies (such as A depends on B, and B depends on C).

[0073] The dependency chain is a specific representation of these paths, and the weight value reflects the importance or risk level of each dependency relationship. By constructing a weighted dependency chain list, the potential impact of each dependency relationship on the overall system can be evaluated more precisely.

[0074] Based on the extracted attribute features, the system constructs a graph structure to represent the dependency relationships between software components. For each pair of dependent components, a weight value is calculated between them, taking into account factors such as the depth of dependence, frequency, and historical vulnerability records, etc. Finally, a list of weighted dependency chains containing all dependency relationships and their weights is generated, providing detailed dependency information for subsequent risk assessment.

[0075] Continuing with the above embodiment, after parsing the dependencies of the requests library, the system discovers that it directly depends on urllib3, and urllib3 indirectly depends on idna. The system constructs a graph structure based on these dependency relationships and assigns weight values to each dependency path. For example, the weight value of requests -> urllib3 is relatively high because this is a common direct dependency; while the weight value of urllib3 -> idna is slightly lower, but still needs to be considered because it is an important indirect dependency. This step lays a solid foundation for subsequent risk assessment.

[0076] 103. Adopt a Bayesian network model combined with a predefined risk index system to calculate the risk scores of the dependency chains in the list of weighted dependency chains, and use an anomaly detection algorithm to identify potential supply chain threat points, so as to generate a risk report including the overall risk score and the supply chain threat points, where the supply chain threat points are the key links or components that can be exploited by attackers to undermine the integrity and security of software during the software development, distribution, and deployment processes;

[0077] The Bayesian network model is a probabilistic graphical model that can capture the conditional dependency relationships between variables and is applicable to complex risk assessment scenarios.

[0078] The risk index system includes multiple dimensions, such as the number of vulnerabilities, update frequency, community activity, etc., to quantify the risk levels of each dependency chain.

[0079] The anomaly detection algorithm is used to identify those data points that significantly deviate from the normal behavior pattern, thereby finding potential supply chain threat points. The risk report summarizes the risk scores and threat point details to help decision-makers comprehensively understand and respond to risks.

[0080] Using the Bayesian network model, combined with a predefined risk index system, the system performs risk scoring on each dependency chain in the list of weighted dependency chains. Then, applying an anomaly detection algorithm (such as Isolation Forest or Autoencoder), it identifies the dependency chains with abnormal behaviors and determines the supply chain threat points. Finally, a detailed risk report is generated, including the overall risk score and threat point profiles, for decision-makers to refer to.

[0081] In the previous requests library case, the system used a Bayesian network model combined with a vulnerability database to evaluate the risk scores of each dependency chain. For example, requests -> urllib3 received a high risk score due to known vulnerabilities. At the same time, the anomaly detection algorithm identified an uncommon dependency of urllib3, pyopenssl, which might be a target for supply chain attacks. The risk report generated by the system not only included these scores but also provided detailed threat point profiles, recommending that the team review the security of pyopenssl and consider alternatives.

[0082] 104. According to the risk report, use a multi-dimensional visualization algorithm to draw a comprehensive risk map to obtain a risk view, and integrate a genetic algorithm and a simulated annealing algorithm to explore adjustment suggestions for high-risk areas of the risk view to generate a risk mitigation strategy; wherein, the comprehensive risk map groups and displays components with similar risk patterns through a hierarchical clustering algorithm, and the risk pattern is the result obtained by judging the similarity of the dependency chain, the supply chain threat point, and the risk score, and the risk mitigation strategy is used to reduce the risk score of the software component.

[0083] The multi-dimensional visualization algorithm is used to transform complex dependency chains and risk scores into an intuitive graphical interface to help users quickly understand the risk distribution. The comprehensive risk map groups and displays components with similar risk patterns through a hierarchical clustering algorithm to facilitate the identification of high-risk areas. The genetic algorithm and the simulated annealing algorithm are used to optimize the adjustment suggestions and explore how to effectively reduce the risk score to generate specific risk mitigation strategies.

[0084] Based on the data in the risk report, the system uses a multi-dimensional visualization algorithm to create an interactive risk map to intuitively display the risk scores and threat points of the dependency chain. Then, through a hierarchical clustering algorithm, the risk patterns are classified to form a grouped display. Next, in combination with a genetic algorithm and a simulated annealing algorithm, the system explores the best adjustment plan for high-risk areas and proposes specific mitigation measures, such as replacing the dependency library or strengthening security audits. Finally, a detailed set of risk mitigation strategies is generated to guide decision-makers to take actions.

[0085] Continuing with the above case, the system generated an interactive risk map showing the risk scores and threat points of the requests library and its dependency libraries. Through the hierarchical clustering algorithm, the map grouped the dependency libraries with similar risk patterns, such as libraries with a high vulnerability rate and low update frequency. Subsequently, the system used the genetic algorithm and the simulated annealing algorithm to propose a number of adjustment suggestions, such as replacing pyopenssl with a more secure alternative or increasing the regular security audits of urllib3. Finally, the system generated a detailed risk mitigation strategy report to guide the development team to take specific measures to reduce the overall risk score.

[0086] Through the implementation of steps 101 to 104, this method realizes the full-process intelligent analysis from parsing software component identification information to generating detailed risk mitigation strategies. First, by accurately parsing the identification information and extracting attribute features, the data accuracy for subsequent analysis is ensured. Second, a weighted dependency relationship chain list is constructed to comprehensively capture direct and indirect dependency relationships, providing a detailed basis for risk assessment. Third, using the Bayesian network model and anomaly detection algorithms, the system can accurately identify potential supply chain threat points and generate reliable risk reports. Finally, through multi-dimensional visualization and optimization algorithms, the system not only intuitively displays the risk distribution but also proposes effective mitigation strategies, significantly enhancing the security and risk management capabilities of the software supply chain. These series of steps complement each other, forming a complete risk management closed-loop, providing strong support for software development, distribution, and deployment.

[0087] To address the complexity and uncertainty in evaluating the risk of dependency chains in the software supply chain, in some embodiments, step 103 adopts a Bayesian network model combined with a pre-defined risk index system to calculate the risk scores of the dependency chains in the weighted dependency relationship chain list, including:

[0088] According to the Bayesian network model, use the Node2Vec or GraphSAGE graph embedding algorithm to vectorize the node paths in the weighted dependency relationship chain list, map the component nodes, edge attributes, and topological structures in the dependency chain to a low-dimensional vector space to obtain dependency vectors; based on the dependency vectors, normalize the vulnerability quantity, update frequency, and community activity indicators through a dynamic weight allocation strategy, where the vulnerability quantity weight is dynamically adjusted according to the severity level of the CVE database, the update frequency weight is calculated with exponential decay in combination with the component version iteration cycle, and the community activity weight is linearly interpolated based on the code submission frequency and issue response time to form a multi-dimensional weighted initial risk score; use the dependency chains with the assigned multi-dimensional weighted initial risk scores, take the risk probability of the parent nodes in the Bayesian network as the feature input of the gradient boosting decision tree, extract the component dependency depth, cross-ecosystem call relationship, and license conflict type as additional features through feature engineering, and use the greedy algorithm to optimize the conditional probability convergence direction when splitting the nodes of the decision tree to obtain the optimized conditional probability; according to the optimized conditional probability, use the Stacking ensemble learning method to perform meta-model fusion on the component vulnerability prediction results of the random forest classifier and the risk propagation coefficient prediction results of the LightGBM regressor, and determine the weights of each base model through cross-validation to output the risk assessment score of the dependency chain.

[0089] In this embodiment, Node2Vec and GraphSAGE are graph embedding algorithms that can convert complex dependency chains into low-dimensional vector forms that are convenient for machine learning models to process. These vector data not only contain the basic connection information between components but also reflect the topological structure characteristics between them, which are used in the subsequent risk assessment process.

[0090] In the embodiments of the present application, first, we use graph embedding technology to map the dependency chain into a low-dimensional space to obtain dependency vectors; second, based on these vectors, we adjust the risk scores of each component through a dynamic weight allocation strategy that takes into account factors such as vulnerability severity, update cycle, and community interaction frequency; third, we input these adjusted risk scores into a gradient boosting decision tree model based on a Bayesian network to optimize the selection of each split point; finally, we integrate the prediction results from the random forest and LightGBM models to obtain a more accurate risk assessment score.

[0091] The following is a specific example:

[0092] Suppose we want to evaluate the security of an open-source project. First, we parse all the dependency declaration files of the project, extract all direct and indirect dependencies from them, and convert them into dependency vectors using the GraphSAGE algorithm; second, we calculate the initial risk scores for each dependency based on the vulnerability severity recorded in the CVE database, the component version iteration cycle, and the time intervals between code commits and issue responses; third, we input these risk scores together with other features extracted from the dependency chain into the gradient boosting decision tree model to optimize the splitting rules for each node; finally, we use the Stacking method to combine the prediction results of the random forest and LightGBM models, and determine the optimal combination weights through cross-validation to obtain the final risk scores for each dependency. Through the above steps, we can not only identify high-risk dependency paths but also propose targeted mitigation measures.

[0093] To further improve the ability to identify potential threat points in the software supply chain, in some embodiments, the use of an anomaly detection algorithm to identify potential supply chain threat points to generate a risk report including an overall risk score and the supply chain threat points, includes:

[0094] Based on the overall risk score, the Isolation Forest algorithm is used to perform unsupervised clustering on the topological feature vectors of the dependency chain. Combining with the characteristics of supply chain attack patterns recorded in the historical attack sample library, the dynamic percentile method based on a sliding window is used to set the anomaly threshold. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of the same type of components in the same period, it is marked as a threat point. Entity recognition is performed on security bulletins through the BERT pre-trained model, and triple information of CVE number, affected version range, and repair solution is extracted by combining dependency syntax analysis. The knowledge graph alignment technology is used to associate and match the extracted threat information with the component metadata in the software bill of materials to generate threat point details including the vulnerability propagation path. All the overall risk scores and the threat point details are summarized and sorted out. The Prophet time series model is used to perform seasonal decomposition on the component vulnerability disclosure frequency and the number of dependency relationship changes. Based on the LSTM neural network, a risk trend prediction model is constructed. The historical risk score, component update interval, and community discussion heat are used as input features to generate a risk fluctuation prediction curve for the next three version cycles. A visualization tool is used to construct a three-dimensional topological map through the ECharts framework to present the component dependency network, use a Sankey diagram to visualize the risk propagation path, and integrate the D3.js force-directed graph to achieve dynamic focus display of high-risk nodes, generating an interactive risk report containing a risk trend overlay layer with a time dimension and a threat point heat distribution layer.

[0095] In this embodiment, the Isolation Forest is an algorithm for anomaly detection, which can effectively identify a small number of outliers from a large amount of normal data. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model that can be used to extract structured information from text, such as CVE numbers and their related repair suggestions. The knowledge graph provides a way to represent the relationships between entities, which helps to associate scattered security intelligence.

[0096] In the embodiment of the present application, first, according to the overall risk assessment result, the Isolation Forest algorithm is used to analyze the topological feature vectors of the dependency chain to determine which components may become threat points; second, we use the BERT model to process public security bulletins, extract important information such as CVE numbers and affected version ranges from them, and associate them with existing component information through knowledge graph technology; third, we summarize all the collected risk scores and detailed threat information, use the Prophet model to analyze time series features, and predict future risk trends with the help of the LSTM network; finally, we use visualization tools, including ECharts and D3.js, to create an interactive risk report interface that not only shows the dependencies between components but also highlights potential high-risk areas.

[0097] The following is a specific example:

[0098] Suppose we want to analyze the security of an open-source software. First, based on its overall risk score, we used the Isolation Forest algorithm to perform unsupervised clustering on its dependency chain and set an anomaly threshold. Second, we applied the BERT model to parse the latest security bulletins, extracted information such as CVE numbers, affected versions, and their repair solutions, and associated and matched them with the component metadata in the software bill of materials through knowledge graph technology. Third, we aggregated all risk scores and threat point details, used the Prophet model to perform seasonal decomposition on the vulnerability disclosure frequency, and built a risk fluctuation prediction model for the next three release cycles based on the LSTM neural network. Finally, we used the ECharts framework to build a three-dimensional topological graph to display the component dependency network, adopted a Sankey diagram to visualize the risk propagation path, and integrated the D3.js force-directed graph to achieve dynamic focusing display of high-risk nodes. Through the above steps, we can not only accurately identify potential threat points in the software supply chain but also provide clear and intuitive guidance on risk mitigation strategies for the development team.

[0099] To further improve the recognition accuracy of the risk propagation path in the software supply chain, in some embodiments, using the dependency chain with the allocated multi-dimensional weighted initial risk scores, taking the risk probability of the parent node in the Bayesian network as the feature input of the gradient boosting decision tree, extracting the component dependency depth, cross-ecosystem call relationship, and license conflict type as additional features through feature engineering, and using the greedy algorithm to optimize the conditional probability convergence direction when splitting nodes of the decision tree to obtain the optimized conditional probability, including:

[0100] Collect and organize the low-dimensional vector representations of the dependency chains and the initial risk scores, extract the component dependency path depth, cross-ecosystem call frequency, and license conflict type as core features, combine the component version iteration interval, vulnerability repair response time, and developer contribution activity to construct an extended feature set, standardize the features to eliminate the dimension difference, and screen out the feature subset that has a significant impact on risk propagation through feature importance analysis to generate a multi-dimensional feature training dataset; based on the Bayesian network model, initialize the network parameters through the correlation analysis of the risk status of the parent node and the dependency strength of the child node, use the non-linear fitting ability of the decision tree model to assist in establishing the initial mapping relationship of the conditional probability between nodes, take the parent node risk probability, component dependency depth, and cross-ecosystem call frequency as the input features of the decision tree, and construct a joint training framework; use the multi-dimensional feature training dataset, adopt an alternating optimization strategy to synchronously update the conditional probability table of the Bayesian network and the splitting rules of the decision tree, calculate the deviation between the model prediction risk probability and the true vulnerability disclosure record through the loss function, combine the backpropagation algorithm to dynamically adjust the network weights and tree structure parameters, and select the optimal splitting node according to the greedy algorithm in each iteration to optimize the convergence direction of the conditional probability until the model converges; according to the joint training results of the joint training framework, weight and correct the conditional probabilities of the child nodes on the high-probability propagation paths in the Bayesian network, refine the probability distribution parameters based on the component version compatibility data in the dependency chain and the statistical results of historical attack events, and feedback the optimized conditional probability values to the Bayesian network to update the dependency relationship strength between nodes to generate optimized conditional probabilities.

[0101] In this embodiment, the Bayesian network is a probabilistic graphical model used to represent the dependency relationships between variables and is suitable for uncertain reasoning. The Gradient Boosting Decision Tree (GBDT) is a powerful machine learning method that is good at dealing with non-linear relationships. The greedy algorithm is an algorithm that makes the best or optimal choice in each step of the selection strategy in the current state, with the expectation that the result is the best or optimal globally. Combining these techniques can effectively analyze and predict the risk propagation paths in the software supply chain.

[0102] In the embodiment of this application, first, we extract key features such as component dependency depth, cross-ecosystem call frequency, and license conflict type from the dependency chain, and construct a comprehensive feature set by combining other extended features; second, we use the Bayesian network model to analyze the relationship between the risk status of the parent node and the dependency strength of the child node to initialize the model parameters of ours; third, we adopt an alternating optimization strategy to synchronously adjust the parameters of the Bayesian network and the decision tree model during each iteration to ensure that the model can accurately reflect the actual situation; finally, based on the results of the joint training framework, we finely adjust the relevant parameters of the high-probability propagation paths in the Bayesian network to improve the accuracy of the overall model.

[0103] The following is a specific example:

[0104] Suppose we want to evaluate the risk propagation of an open-source software project. First, we collect the version information, license types, and historical vulnerability records of all dependent components of the project, calculate the initial risk scores of each component, and extract core features such as component dependency depth, cross-ecosystem call frequency, and license conflict types. Second, based on the Bayesian network model, we analyze the tightness of the dependency relationships between components and convert these relationships into the initial parameters of the model. Third, we use the training dataset containing the above features and continuously adjust the parameters of the Bayesian network and decision tree model through an alternating optimization strategy until the model prediction results match the actual vulnerability disclosure records. Finally, according to the feedback of the joint training framework, we optimize and adjust the conditional probabilities of the nodes involved in the high-risk propagation paths in the Bayesian network, so as to more accurately reflect the changes in the dependency strength between components. Through the above steps, we not only improve the understanding of the risk propagation paths in the software supply chain but also provide strong support for subsequent risk management.

[0105] In order to further improve the recognition accuracy of potential threat points in the software supply chain, in some embodiments, based on the overall risk score, the isolation forest algorithm is used to perform unsupervised clustering on the topological feature vectors of the dependency relationship chain. Combining the supply chain attack pattern features recorded in the historical attack sample library, the dynamic percentile method based on a sliding window is used to set the anomaly threshold. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of the same type of components in the same period, it is marked as a threat point, including:

[0106] According to the overall risk score, extract the topological feature vectors of the dependency chain, including component node degree centrality, dependency path length, and cross-ecosystem call density. Combine the characteristics of supply chain attack patterns recorded in the historical attack sample library to construct a multi-dimensional risk feature space. Normalize the features to eliminate the dimension difference, and reduce the dimension through principal component analysis to improve the calculation efficiency, generating an anomaly detection model adapted to the current software ecosystem. Use an unsupervised clustering algorithm to group the topological feature vectors, and combine the labeling of high-risk samples in the historical attack sample library to train the anomaly detection model. Optimize the algorithm parameters through model evaluation metrics to generate a detection model that can identify outliers in the multi-dimensional risk feature space. Based on the historical data of the dependency chain, statistically analyze the risk score distribution interval of similar components in the normal operation state, combine the high-risk sample set labeled by domain experts to set the initial anomaly determination boundary, and use the dynamic percentile method based on a sliding window to monitor the vulnerability disclosure frequency in the component warehouse update log in real time. Dynamically expand the anomaly feature dimension according to the characteristics of newly emerging attack vectors to generate a dynamic threshold that matches the current threat environment. Use the anomaly detection model combined with the dynamic threshold to analyze the risk score mutation phenomenon during the component version iteration in the dependency chain, identify complex threat points that simultaneously meet the topological structure anomaly, vulnerability density exceeding the standard, and community repair response lag. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of similar components in the same period, mark it as a supply chain threat point and generate a list of supply chain threat points.

[0107] In this embodiment, Isolation Forest is an effective unsupervised learning method for identifying a small number of outliers from data. The topological feature vectors include information such as component node degree centrality, dependency path length, and cross-ecosystem call density, which are used to evaluate the complex relationships and potential risks between components. Principal Component Analysis (PCA) is a dimensionality reduction technique that helps simplify high-dimensional datasets while retaining as much important information from the original data as possible. The dynamic percentile method allows the threshold for anomaly detection to be adjusted dynamically based on the latest data, thus better adapting to the changing security environment.

[0108] In the embodiments of the present application, first, we extracted the key topological features of the dependency chain based on the overall risk score, and constructed a multi-dimensional risk feature space through normalization processing and principal component analysis; second, we used the Isolation Forest algorithm to perform unsupervised clustering on these feature vectors, and optimized the model parameters in combination with historical attack samples to ensure effective identification of outliers; third, based on the historical data of the dependency chain and the high-risk sample set provided by domain experts, a preliminary anomaly determination boundary was set, and the anomaly threshold was updated in real time using the dynamic percentile method based on a sliding window; finally, we applied the above model to the actual software supply chain environment, and identified those complex threat points with abnormal topological structures, high vulnerability densities, and lagging community repair responses by analyzing the mutation of risk scores during component version iteration.

[0109] The following is a specific example:

[0110] Suppose we want to conduct a security assessment on an open-source software project. First, we extracted features such as the node degree centrality, dependency path length, and cross-ecosystem call density of all dependent components based on its overall risk score, and generated a multi-dimensional risk feature space adapted to the current project through PCA processing; second, we applied the Isolation Forest algorithm to perform unsupervised clustering on these features, optimized the model parameters in combination with historical attack cases, and established an efficient anomaly detection model; third, based on the historical data of the project and the opinions of domain experts, an initial anomaly determination boundary was set, and the anomaly threshold was continuously monitored and adjusted through the dynamic percentile method; finally, we used this anomaly detection model to analyze the change in risk scores during component version updates, and found that the vulnerability density of a component was significantly higher than three times the standard deviation of similar components, and there was also a problem of lagging repair response, so it was marked as a supply chain threat point. Through the above steps, we not only successfully identified potential security threats, but also provided an important basis for subsequent risk management and mitigation strategies.

[0111] In order to further improve the visualization analysis ability of software supply chain risks, in some embodiments, according to the risk report in step 104, a multi-dimensional visualization algorithm is used to draw a comprehensive risk map to obtain a risk view, including:

[0112] According to the risk report, integrate the dependency chain topology data, threat point spatial location information, and risk score time series to construct a structured data set containing component coordinates, risk level labels, and threat propagation direction vectors; use multi-dimensional visualization algorithms to map the dependency relationship chain into connection lines in three-dimensional space, convert the risk score into a node size parameter, and label the threat point position as a pulse signal marker to generate an interactive spatial distribution map; based on the visual representation, use density-aware clustering technology to automatically demarcate the high-risk component aggregation area, generate risk pattern groups according to the development team, license type, and vulnerability history similarity of the components, and achieve a hierarchical display of the threat propagation path; combine the network topology structure, design a circular layout algorithm to optimize the node arrangement, use gradient color coding to represent the risk trend changes in different time windows, integrate the focus + context visualization technology to achieve a scalable display of large-scale dependency networks, and generate a comprehensive risk map containing spatio-temporal multi-dimensional features; display the vulnerability infection intensity between components through heat layer superposition, use a dynamic streamline diagram to present the diffusion path of risks in the supply chain, and provide a risk node penetration query function to reveal the underlying threat evidence chain to form an actionable risk view.

[0113] In this embodiment, the multi-dimensional visualization algorithm refers to a series of technologies for presenting complex data relationships, such as three-dimensional space mapping, heat layer superposition, etc. Density-aware clustering is an algorithm that can identify dense areas in a data set and helps to discover groups of high-risk components. The circular layout algorithm is a method specifically designed to optimize the node arrangement, making the network diagram clearer and easier to read. These technologies and methods work together to effectively help users understand complex software supply chain risks.

[0114] In the embodiment of this application, first, we integrated all relevant data extracted from the risk report, including the topological structure of the dependency chain, the spatial locations of each threat point, and the risk scores changing over time, to form a detailed structured data set; second, we mapped this data into three-dimensional space, using connection lines to represent the dependency relationship chain, the node size reflecting the risk score, and the pulse signal marker indicating the position of the threat point; third, we used density-aware clustering technology to identify the areas where high-risk components gather and classified these areas according to factors such as the development team, license type, and vulnerability history; finally, we applied the circular layout algorithm to optimize the visual effect of the entire network diagram, used color gradient coding to display the risk trends in different periods, and implemented an interactive large-scale network browsing function, thus generating a comprehensive risk map that is both detailed and easy to understand.

[0115] The following is a specific example:

[0116] Suppose we want to create a risk view for a large open-source project. First, we collect the topological data of all the dependent components of the project, the spatial locations of the threat points, and their risk scores over time, and construct a dataset containing component coordinates, risk level labels, and threat propagation direction vectors. Second, we use a multi-dimensional visualization tool to convert the dependency chains into connection lines in three-dimensional space, adjust the node sizes to visually display the risk scores, and use pulse signals to mark and highlight the threat points. Third, we apply density-aware clustering techniques to identify high-risk component clusters and classify these areas according to development team affiliation, license type, and historical vulnerability situations. Finally, we adopt a circular layout to optimize the node distribution, use color gradients to show the risk trend changes over the past few months, and integrate the focus+context technique to allow users to zoom in and view the detailed information of specific components, while providing a risk node penetration query function to reveal the potential threat evidence chain. Through the above steps, we can not only clearly display the overall risk situation of the software supply chain but also provide the ability to deeply explore each risk point, greatly enhancing the effectiveness and efficiency of risk management.

[0117] In order to further improve the accuracy and security of dependency management in the software supply chain, in some embodiments, the positioning of the node paths associated between the software components according to the attribute characteristics of the software components to generate the dependency chains in the node paths and assigning weight values to the dependency chains to generate a weighted dependency chain list includes:

[0118] Using the parsed software component identification information submitted by the user to extract structured attribute characteristics including version numbers, build tool types, and dependency declaration file formats; establishing a multi-level index model of component dependencies according to the attribute characteristics of the software components, identifying explicit declared dependencies and implicit deduced dependencies through a dependency resolution engine, and generating a dependency network containing version constraint conditions; based on the dependency network, tracing the call records of third-party libraries dynamically loaded during the build process, analyzing the interface call paths covered by test cases, and verifying the actually effective dependency relationships in combination with continuous integration logs to generate dependency chain records accurate to specific call links; classifying the dependency chain records, dividing direct and indirect dependencies according to the dimensions of whether the dependency introduction method is necessary for the development environment and whether it is forcibly loaded at runtime, and assigning exponentially increasing propagation weights to deeply nested cross-ecosystem dependencies to generate preliminary weighted dependency chains; based on the risk report, establishing a weight correction rule engine, dynamically adjusting the weight coefficients according to the security factors such as the reputation rating of component maintainers, the binary file signature verification status, and the integrity of the dependency lock file, and verifying the rationality of weight assignment through dependency propagation simulation, and finally generating a weighted dependency chain list optimized through multiple rounds of iteration.

[0119] In this embodiment, the dependency resolution engine refers to a technology used to identify and analyze the dependency relationships between components in a software project. The multi-level index model is a data structure that can effectively organize and retrieve complex dependency relationships. The propagation weight is used to quantify the influence degree of the dependency relationship in the network, helping to distinguish the importance of direct dependencies and indirect dependencies.

[0120] In the embodiment of the present application, first, we extract key attributes such as version numbers and build tool types from the software component identification information provided by the user to understand the specific situation of each component; second, we use this information to construct a detailed dependency relationship network, including both explicitly declared dependencies and implicit dependencies; third, we deeply explore the call records and test cases during the construction process to ensure that all actually effective dependencies are accurately recorded; finally, we classify these dependency chains according to their nature and assign appropriate weights according to their positions and roles in the network, thereby generating a comprehensive and detailed list of weighted dependency chains.

[0121] The following is a specific example:

[0122] Suppose we want to create a list of dependency chains for a mobile application. First, we parse all the component identification information of the project and extract detailed attributes including version numbers and build tool types; second, based on this information, we establish a multi-level index model of component dependencies, use the dependency resolution engine to identify all explicit and implicit dependencies, and form a detailed dependency relationship network; third, we carefully trace all the call records of third-party libraries dynamically loaded during the construction process, analyze the interface call paths covered by the test cases, and verify the actually effective dependency relationships through the continuous integration logs to obtain a record of dependency chains accurate to specific call links; then, we classify the dependencies into direct and indirect categories according to the criterion of whether they are necessary for the development environment or only loaded at runtime, and assign exponentially increasing propagation weights to those dependencies deeply nested across ecosystems; finally, we adjust the weight coefficients of the dependency chains with reference to factors such as the credibility rating of the component maintainer and the binary file signature status, and verify the rationality of the weight assignment by simulating the propagation of dependency relationships. Through the above steps, we not only clarify the dependency relationships between components, but also evaluate their importance and potential risks in the entire software supply chain, greatly improving the maintainability and security of the project.

[0123] The present application considers that in order to accurately evaluate the risk level in the dependency chain, a comprehensive set of risk assessment formula systems has been developed; this system is based on the low-dimensional vector representation of the dependency chain (dependency vector ), combined with multi-dimensional risk indicators, and by introducing a time decay factor and non-linear transformation, quantifies each risk attribute to generate a quantified risk characteristic ; To further improve accuracy, the weights are dynamically adjusted according to the importance of different risk indicators, and an adaptive learning rate and Bayesian optimization method are introduced to generate risk weights reflecting the influence of different risk factors ; Finally, the adjusted weights are used to calculate the weighted quantization of risk characteristics, and an exponential decay and non-linear term are introduced to comprehensively evaluate the overall risk level of the dependency chain, obtaining an initial risk score , so a new alternative solution is proposed, which includes:

[0124] The initial risk score is calculated based on the dependency vector using a predefined risk indicator system in combination with a multi-dimensional weight adjustment mechanism, including:

[0125] The data in the dependency chain is converted into a low-dimensional vector representation using a graph embedding algorithm, and a regularization term is introduced to prevent overfitting, generating a dependency vector ; Among them, the dependency vector is obtained by calculating the following formula:

[0126] ;

[0127] Wherein, is the dependency vector of the dependency chain, is the graph data corresponding to the dependency chain, is the time since the last update, is the mean of the time intervals, is the standard deviation of the time intervals, are the weight matrix and bias vector of the first layer respectively, is the activation function, is the number of network layers, represents the set of all parameters, is the th parameter, is the weight matrix of the th layer in the neural network, is the bias vector of the th layer in the neural network, is the standard deviation of the time intervals;

[0128] The following provides a detailed explanation of each parameter:

[0129] : The low-dimensional vector representation of the dependency chain; The complex graph data in the dependency chain is converted into a low-dimensional vector through a graph embedding algorithm to capture the key features of the dependency chain;

[0130] : Graph data (adjacency matrix or edge list) corresponding to the dependency chain; extracted from the actual dependency network, describing the connection relationships between nodes;

[0131] : The weight matrix and bias vector of the neural network at the

[0132] layer; optimized through the training process (such as gradient descent method) to capture non-linear patterns in the graph data;

[0133] : The number of layers of the neural network; the multi-layer structure can learn and abstract features in the graph data more deeply;

[0134] : The time since the last update; records the most recent update time of each dependency;

[0135] : The mean of the time intervals, which statistically calculates the average of the update time intervals of all dependencies.

[0136] : The standard deviation of the time intervals, which statistically calculates the standard deviation of the update time intervals of all dependencies.

[0137] : The set of all parameters, including the weight matrix and bias vector, containing all the parameters optimized during the model training process.

[0138] : The th parameter, obtained through the model training process, specifically referring to the elements in each weight matrix and bias vector.

[0139] The design reasons for each item are introduced as follows:

[0140] Deep neural network : Use a multi-layer neural network to encode the graph data of the dependency chain. Through layer-by-layer transformation and non-linear activation functions, capture complex dependency relationships and features. The weight matrix and bias vector of each layer are optimized through training to ensure that the model can effectively learn and express the key information in the graph data.

[0141] Time decay factor: Introduce a time decay factor in the form of a Gaussian distribution, so that newer dependencies have a greater impact on the final result, while the impact of older dependencies gradually weakens. This reflects the time sensitivity and ensures that the model can dynamically adapt to changes in dependencies.

[0142] Regularization term : By summing the squares of all parameters, introduce an L2 regularization term to prevent the model from overfitting. The regularization term helps to maintain the generalization ability of the model and ensures that it can also perform well on unseen data.

[0143] Adding the neural network output to the time decay factor and the regularization term is to comprehensively consider the static characteristics of the dependency chain (encoded by the neural network) and the time-varying factors (through the time decay factor), while adding regularization to prevent overfitting. This combination enables the model to comprehensively capture the spatio-temporal characteristics of the dependency chain and maintain good generalization performance.

[0144] The following is a specific example:

[0145] Suppose there is a simple dependency chain graph data , which contains 5 nodes and 7 edges; use a two-layer neural network , with the ReLU activation function , and introduce a regularization term to prevent overfitting;

[0146] The specific parameters are as follows:

[0147] : The weight matrix and bias vector of the first layer;

[0148] : The weight matrix and bias vector of the second layer;

[0149] ;

[0150] contains all parameters. Suppose , and each parameter ;

[0151] Substitute into the formula to get:

[0152] ;

[0153] Suppose is known and has been trained, and the specific value of can be calculated:

[0154] ;

[0155] Conclusion explanation:

[0156] Dependency vector Provides a compact representation of the dependency chain, capturing the importance and interdependencies of individual nodes. Lower values indicate relatively less risk for a node, while higher values imply greater potential risk. This result provides a solid foundation for subsequent risk assessment, helping decision-makers better understand and address potential risks in the dependency chain, especially in areas such as supply chain management and cybersecurity.

[0157] According to the said dependency vector , combined with the risk index system, quantify the risk attributes of the dependency chain, and introduce a time decay factor and non-linear transformation to obtain the quantified risk characteristics; where the said quantified risk characteristics Are obtained by calculating the following formula:

[0158] ;

[0159] Wherein, Is the quantified risk characteristic, Is the number of risk indicators, Is the Th weight of the Th risk indicator, Is a non-linear transformation function used to increase the expressive power, Represents the vector dot product, Are respectively And The L2 norms of;

[0160] The following gives a detailed explanation of each parameter:

[0161] : The Th quantified risk characteristic of the Th dependency chain; it synthesizes the influence of multiple risk indicators, and quantifies the risk attributes of the dependency chain by introducing a time decay factor and non-linear transformation; this value reflects the overall risk level of a specific dependency chain in the

[0162] Th risk dimension, and is an important basis for subsequent risk assessment and decision-making; : The number of risk indicators; each risk indicator

[0163] : Is the The weights of individual risk indicators; these weights reflect the relative importance of different risk indicators in the overall risk assessment; by dynamically adjusting the importance of each risk indicator, the model can better adapt to different application scenarios and ensure the accuracy and reliability of the risk assessment results; the selection of weights can be optimized based on historical data or expert knowledge;

[0164] : is a non-linear transformation function used to increase the expressive power and capture complex risk patterns; common choices include the hyperbolic tangent function (tanh), the Sigmoid function, etc.; the non-linear transformation enables the model to better handle complex non-linear relationships, improving the flexibility and accuracy of risk assessment; by introducing , the model can adjust the expression of risk characteristics over a wider range, enhancing its adaptability and robustness;

[0165] : represents the dependence vector and the th risk indicator's feature vector ; the dot product measures the similarity in direction between two vectors, and the larger the value, the closer the two vectors are; by calculating the dot product, the degree of association between the dependence relationship chain and a specific risk indicator can be evaluated, providing a basis for risk quantification;

[0166] : is the feature vector related to the th risk indicator ; it transforms the risk indicator into a representation in a low-dimensional space, facilitating mathematical operations and model processing; each feature vector contains multiple dimensions, and each dimension represents a feature or attribute, reflecting the specific information of the risk indicator; through the feature vector, the model can more precisely capture and quantify the impact of the risk indicator on the dependence relationship chain;

[0167] : are the L2 norms (i.e., Euclidean norms) of the dependence vector and the feature vector respectively; the L2 norm measures the length or magnitude of a vector and is used to standardize the dot product calculation to ensure the effectiveness of the cosine similarity; by dividing by the L2 norm, the influence of the vector length can be eliminated, focusing on their direction similarity, thus more accurately evaluating the correlation between the dependence relationship chain and the risk indicator.

[0168] The following is a specific example:

[0169] Suppose there are three risk indicators , and the corresponding feature vectors are

[0170] ;

[0171] The weights are respectively ;

[0172] Assume , and the calculation gives:

[0173] ;

[0174] Assume the calculation results are respectively:

[0175] ;

[0176] The specific numerical value of the quantified risk characteristic is obtained. This numerical value reflects the risk level of the dependency chain in the th risk dimension. A lower quantified risk characteristic value indicates that the risk of the dependency chain in this risk dimension is relatively small and the system is relatively stable; a higher quantified risk characteristic value means that there are greater potential risks and further attention and measures are needed to reduce the risks.

[0177] Dynamically adjust the weights of each index according to the importance of different risk indicators, and introduce an adaptive learning rate and Bayesian optimization method to reflect the influence degree of different risk factors on the overall risk, and generate risk weights; among them, the said risk weights are calculated through the following formula:

[0178] ;

[0179] Among them, is the risk weight, used to measure the difference between the predicted value and the actual value, is the mean value of the loss function, is the standard deviation of the loss function, is the initial weight vector, is the optimal weight configuration obtained through Bayesian optimization, is the number of samples, is the th predicted value of the sample, is the th true label or target value of the sample, output by the model, is the square of the prediction error of each sample, is the weight matrix;

[0180] The following gives a detailed explanation of each parameter:

[0181] : The risk weight vector after dynamic adjustment; it optimizes the weight configuration of each risk indicator according to the historical performance of the loss function and current data, enabling the model to more accurately reflect the impact degree of different risk factors on the overall risk; by introducing an adaptive learning rate and Bayesian optimization method, the model can flexibly adjust the weights to improve the accuracy and robustness of risk assessment;

[0182] : The mean of the loss function, representing the historical average level of the difference between the model's predicted value and the actual value; by comparing the current loss with the historical mean, the model can evaluate its own performance changes and adjust the weights accordingly to optimize future performance;

[0183] : The standard deviation of the loss function, measuring the variation range of the loss values; a larger standard deviation indicates greater loss fluctuations and the model may be less stable; a smaller standard deviation means the losses are more consistent and the model performs more stably; by introducing the standard deviation, the model can respond more sensitively to loss changes and adjust the weight configuration in a timely manner;

[0184] : The initial weight vector of the model; it is the starting point for weight adjustment, usually based on random initialization or the results of pre-training; through gradual optimization, the model can start from the initial weight vector to find the optimal weight configuration to minimize the loss function;

[0185] : The optimal weight configuration obtained through Bayesian optimization or other optimization methods; it represents the weight combination that minimizes the loss function on the given dataset; through continuous iterative optimization, the model can gradually approach this optimal configuration to improve the prediction accuracy;

[0186] : The total number of samples in the dataset; it is used to calculate the average loss to ensure that the loss function reflects the performance of the entire dataset rather than the anomalies of individual samples; a larger number of samples can make the loss estimate more stable and reliable;

[0187] : is the predicted value of the

[0188] : is the true label or target value of the

[0189] : is the square of the prediction error of the th sample; it amplifies the impact of larger errors, making the model pay more attention to those samples with larger prediction deviations; by minimizing the sum of the squares of the prediction errors, the model can optimize its own parameters and improve the accuracy of prediction;

[0190] : is the weight matrix in the model, used to connect the neurons in the input layer and the hidden layer, or between the hidden layers; it determines the learning ability and expressive ability of the model; through optimization algorithms such as the gradient descent method, the model can continuously adjust the weight matrix to minimize the loss function and improve the prediction performance.

[0191] The following is a specific example:

[0192] Assume the mean of the loss function , the standard deviation , the initial weight , the optimal weight configuration , the number of samples , the sum of the squares of the prediction errors ;

[0193] Substituting the values gives:

[0194] ;

[0195] Assume the gradient , calculated as:

[0196] ;

[0197] The risk weights are obtained, and these weights reflect the relative importance of different risk indicators in the overall risk assessment. A lower risk weight indicates that the risk indicator has a smaller impact on the overall risk and the system is relatively stable; a higher risk weight means there is a greater potential risk and further attention and measures are needed to reduce the risk.

[0198] Using the said risk weights to perform weighted calculations on the quantified risk characteristics , and introducing exponential decay and non-linear terms to increase the expressive ability of the model, comprehensively evaluating the risk level of the said dependency chain, and obtaining the initial risk score ; where, the said initial risk score is calculated through the following formula:

[0199] ;

[0200] Where, is the initial risk score, is the number of risk indicators, is the adjusted weight of the th risk indicator, and are the mean and standard deviation of the quantified risk characteristics respectively, is the shortest path distance between nodes in the dependency chain, and

[0201] are the mean and standard deviation of the shortest path distance respectively.

[0202] : The initial risk score of the dependency chain, which synthesizes the influence of multiple risk indicators and their weights; it increases the expressive power of the model by introducing exponential decay and non-linear terms, reflecting the overall risk level of the dependency chain; a lower initial risk score indicates relatively less risk for this dependency chain and a more stable system; a higher score means there are greater potential risks, requiring further attention and measures to reduce risks;

[0203] : Represents the number of risk indicators; each risk indicator is a measure of the risk of a specific aspect of the dependency chain, such as the number of vulnerabilities, update frequency, vendor reputation, etc.; by considering multiple risk indicators, the overall risk situation of the dependency chain can be evaluated more comprehensively; The size of

[0204] depends on the number of risk factors included in the specific application scenario; : Is the adjusted weight of the

[0205] th risk indicator; these weights reflect the relative importance of different risk indicators in the overall risk assessment; by dynamically adjusting the importance of each risk indicator, the model can better adapt to different application scenarios, ensuring the accuracy and reliability of the risk assessment results; the selection of weights can be optimized based on historical data or expert knowledge; is the mean of all quantified risk characteristics indicating the average level of risk characteristics; is the standard deviation of the quantified risk characteristics, measuring the variation range of the risk characteristic values; by introducing the mean and standard deviation, the model can standardize the quantified risk characteristics, eliminate the influence of outliers, and enhance the sensitivity to risk changes; this helps to more accurately capture the distribution of risk characteristics, thereby improving the accuracy of risk assessment;

[0206] : is the shortest path distance between nodes in the dependency chain; it measures the degree of connection tightness between nodes in the dependency chain; a shorter distance usually means a stronger dependency relationship and a higher possibility of risk propagation, while a longer distance may indicate a weaker dependency relationship and a smaller possibility of risk propagation; by considering the shortest path distance, the model can more accurately evaluate the structural complexity and potential risks of the dependency chain;

[0207] and : is the mean of all shortest path distances and represents the average connection tightness between nodes in the dependency chain; is the standard deviation of the shortest path distance, which measures the variation range of the path distance; by introducing the mean and standard deviation, the model can standardize the shortest path distance, eliminate the influence of outliers, and enhance the sensitivity to path changes; this helps to more precisely capture the structural characteristics of the dependency chain, thereby improving the accuracy of risk assessment.

[0208] The following is a specific example:

[0209] Suppose the mean of the quantified risk characteristics, the standard deviation , the mean of the shortest path distance, the standard deviation , and the specific shortest path distance ;

[0210] Substituting the specific values into the above formula for calculation gives:

[0211] ;

[0212] The calculation result is:

[0213] ;

[0214] Through the above calculation, the initial risk score is obtained, and this score reflects the overall risk level of the dependency chain. A lower initial risk score indicates that the risk of this dependency chain is relatively small and the system is relatively stable; a higher score means there are greater potential risks and further attention and measures are needed to reduce the risks. This evaluation result can provide strong support for fields such as supply chain management and network security, helping decision-makers better understand and respond to potential risks.

[0215] Figure 2 This figure shows the structural schematic diagram of a multi-source software supply chain intelligent analysis device (or system) provided by an embodiment of the present application. As Figure 2 shown, the device includes:

[0216] A parsing module 21, configured to parse the identification information of the software components submitted by the user, and extract the attribute features of the software components from the identification information;

[0217] A positioning module 22, configured to locate the node paths associated between the software components according to the attribute features of the software components, generate a dependency relationship chain in the node paths, assign weight values to the dependency relationship chain, and generate a weighted dependency relationship chain list, wherein the dependency relationship chain includes a direct dependency relationship chain and an indirect dependency relationship chain;

[0218] A calculation module 23, configured to calculate the risk scores of the dependency relationship chains in the weighted dependency relationship chain list by using a Bayesian network model in combination with a predefined risk index system, and identify potential supply chain threat points by using an anomaly detection algorithm, so as to generate a risk report including an overall risk score and the supply chain threat points, wherein the supply chain threat points are key links or components that can be exploited by attackers to damage the integrity and security of the software during the software development, distribution, and deployment processes;

[0219] A drawing module 24, configured to draw a comprehensive risk map by using a multi-dimensional visualization algorithm according to the risk report, obtain a risk view, and fuse a genetic algorithm and a simulated annealing algorithm to explore adjustment suggestions for high-risk areas of the risk view, so as to generate a risk mitigation strategy; wherein the comprehensive risk map is grouped and displayed by a hierarchical clustering algorithm to show components with similar risk patterns, and the risk pattern is a result obtained by judging the similarity of the dependency relationship chain, the supply chain threat points, and the risk scores, and the risk mitigation strategy is used to reduce the risk scores of the software components.

[0220] Figure 2 The multi-source software supply chain intelligent analysis device described above can execute Figure 1 The multi-source software supply chain intelligent analysis method described in the embodiments shown, and its implementation principle and technical effects will not be elaborated. For the multi-source software supply chain intelligent analysis device in the above embodiments, the specific manners in which each module and unit execute operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0221] In a possible design, Figure 2 The multi-source software supply chain intelligent analysis device in the embodiments shown can be implemented as a computing device, such as Figure 3 shown, and the computing device can include a storage component 31 and a processing component 32;

[0222] The storage component 31 stores one or more computer instructions, and one or more of the computer instructions are called and executed by the processing component 32.

[0223] The processing component 32 is used for the above-mentioned Figure 1 multi-source software supply chain intelligent analysis method of the above-mentioned embodiment.

[0224] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0225] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0226] Of course, the computing device may necessarily further include other components, such as input / output interfaces, display components, communication components, etc.

[0227] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above-mentioned peripheral interface module may be an output device, an input device, etc.

[0228] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0229] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above-mentioned processing component, storage component, etc. may be basic server resources leased or purchased from a cloud computing platform.

[0230] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 multi-source software supply chain intelligent analysis method of the above-mentioned embodiment.

[0231] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0232] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0233] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A multi-source software supply chain intelligent analysis method, characterized in that: include: Parsing identification information of a software component submitted by a user, and extracting attribute features of the software component from the identification information; According to the attribute characteristics of the software components, the node paths associated between the software components are located to generate a dependency chain in the node path, and a weight value is assigned to the dependency chain to generate a weighted dependency chain list, wherein the dependency chain includes a direct dependency chain and an indirect dependency chain; A Bayesian network model is used in combination with a predefined risk indicator system to calculate the risk score of the dependency chain in the weighted dependency chain list, and an anomaly detection algorithm is used to identify potential supply chain threat points to generate a risk report containing an overall risk score and the supply chain threat points, wherein the supply chain threat points are key links or components that are used by attackers to undermine the integrity and security of software during software development, distribution and deployment; According to the risk report, a comprehensive risk map is drawn using a multidimensional visualization algorithm to obtain a risk view, and a genetic algorithm and a simulated annealing algorithm are integrated to explore adjustment suggestions for high-risk areas of the risk view to generate a risk mitigation strategy; wherein the comprehensive risk map is grouped and displayed using a hierarchical clustering algorithm to display components with similar risk patterns, wherein the risk pattern is obtained by similarity judgment of the dependency chain, the supply chain threat point and the risk score, and the risk mitigation strategy is used to reduce the risk score of the software component; The Bayesian network model is used in combination with a predefined risk indicator system to calculate the risk score of the dependency chain in the weighted dependency chain list, including: According to the Bayesian network model, the node paths in the weighted dependency chain list are vectorized using the Node2Vec or GraphSAGE graph embedding algorithm, and the component nodes, edge attributes and topological structures in the dependency chain are mapped to a low-dimensional vector space to obtain a dependency vector; Based on the dependency vector, the number of vulnerabilities, update frequency, and community activity indicators are normalized through a dynamic weight allocation strategy, where the weight of the number of vulnerabilities is dynamically adjusted according to the severity level of the CVE database, the update frequency weight is calculated by exponential decay in combination with the component version iteration cycle, and the community activity weight is generated based on linear interpolation of code submission frequency and issue response time to form a multi-dimensional weighted initial risk score; Using the dependency chain with assigned multi-dimensional weighted initial risk scores, the parent node risk probability in the Bayesian network is used as the feature input of the gradient boosting decision tree. Through feature engineering, component dependency depth, cross-ecological call relationship and license conflict type are extracted as additional features. The greedy algorithm is used to optimize the conditional probability convergence direction when the decision tree splits the node, and the optimized conditional probability is obtained. According to the optimized conditional probability, the Stacking ensemble learning method is used to perform meta-model fusion on the component vulnerability prediction results of the random forest classifier and the risk propagation coefficient prediction results of the LightGBM regressor. The weights of each base model are determined through cross-validation, and the risk assessment score of the dependency chain is output.

2. The method according to claim 1, characterized in that The use of anomaly detection algorithms to identify potential supply chain threat points to generate a risk report containing an overall risk score and the supply chain threat points includes: Based on the overall risk score, the isolation forest algorithm is used to perform unsupervised clustering on the topological feature vectors of the dependency chain. Combined with the supply chain attack pattern characteristics recorded in the historical attack sample library, the dynamic percentile method based on the sliding window is used to set the abnormal threshold. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of similar components in the same period, it is marked as a threat point; The BERT pre-trained model is used to identify entities in security bulletins, and dependency syntactic analysis is used to extract triple information of CVE number, affected version range, and repair solution. The knowledge graph alignment technology is used to associate and match the extracted threat information with the component metadata in the software bill of materials, generating threat point details including the vulnerability propagation path. All the overall risk scores and threat point details are summarized and collated, and the Prophet time series model is used to seasonally decompose the component vulnerability disclosure frequency and the number of dependency changes. A risk trend prediction model is built based on the LSTM neural network, and the historical risk scores, component update intervals, and community discussion heat are used as input features to generate risk fluctuation prediction curves for the next three version cycles; Visualization tools are used to build a three-dimensional topological map through the ECharts framework to present the component dependency network, a Sankey diagram is used to visualize the risk propagation path, and D3.js force-directed graph is integrated to achieve dynamic focus display of high-risk nodes, generating an interactive risk report that includes a time dimension risk trend overlay layer and a threat point thermal distribution layer.

3. The method according to claim 1, characterized in that The method utilizes the dependency chain with assigned multi-dimensional weighted initial risk scores, uses the parent node risk probability in the Bayesian network as the feature input of the gradient boosting decision tree, extracts component dependency depth, cross-ecological call relationship and license conflict type as additional features through feature engineering, and uses a greedy algorithm to optimize the conditional probability convergence direction when the decision tree splits the node, and obtains the optimized conditional probability, including: Collect and organize the low-dimensional vector representation of the dependency chain and the initial risk score, extract component dependency path depth, cross-ecosystem call frequency, and license conflict type as core features, build an extended feature set based on component version iteration interval, vulnerability repair response time, and developer contribution activity, standardize the features to eliminate dimensional differences, and screen out feature subsets that have a significant impact on risk propagation through feature importance analysis to generate a multi-dimensional feature training data set; Based on the Bayesian network model, the network parameters are initialized by analyzing the correlation between the risk status of the parent node and the dependency strength of the child node. The nonlinear fitting ability of the decision tree model is used to assist in establishing the initial mapping relationship between the conditional probabilities of the nodes. The parent node risk probability, component dependency depth, and cross-ecological call frequency are used as input features of the decision tree to build a joint training framework. Using the multi-dimensional feature training data set, an alternating optimization strategy is adopted to synchronously update the conditional probability table of the Bayesian network and the splitting rule of the decision tree, the deviation between the risk probability predicted by the model and the actual vulnerability disclosure record is calculated through the loss function, and the network weight and tree structure parameters are dynamically adjusted in combination with the back-propagation algorithm. In each iteration, the optimal splitting node is selected according to the greedy algorithm to optimize the convergence direction of the conditional probability until the model converges; According to the joint training results of the joint training framework, the conditional probabilities of the sub-nodes of the high-probability propagation paths in the Bayesian network are weighted and modified, and the probability distribution parameters are refined based on the component version compatibility data in the dependency chain and the statistical results of historical attack events. The optimized conditional probability values ​​are fed back to the Bayesian network to update the dependency strength between nodes and generate optimized conditional probabilities.

4. The method according to claim 2, characterized in that: Based on the overall risk score, the isolation forest algorithm is used to perform unsupervised clustering on the topological feature vectors of the dependency chain. Combined with the supply chain attack pattern characteristics recorded in the historical attack sample library, the dynamic percentile method based on the sliding window is used to set the abnormal threshold. When the vulnerability density of the component dependency path exceeds 3 times the standard deviation of similar components in the same period, it is marked as a threat point, including: According to the overall risk score, extract the topological feature vector of the dependency chain, including component node degree centrality, dependency path length, and cross-ecological call density, and combine the supply chain attack pattern characteristics recorded in the historical attack sample library to construct a multidimensional risk feature space, normalize the features to eliminate dimensional differences, and reduce the dimension through principal component analysis to improve computing efficiency, and generate an anomaly detection model adapted to the current software ecosystem; An unsupervised clustering algorithm is used to group topological feature vectors. The high-risk sample annotations in the historical attack sample library are combined to train an anomaly detection model. The algorithm parameters are optimized through model evaluation indicators to generate a detection model that can identify outliers in the multi-dimensional risk feature space. Based on the historical data of the dependency chain, the risk score distribution range of similar components under normal operation is counted, and the initial abnormality judgment boundary is set in combination with the high-risk sample set annotated by domain experts. The dynamic percentile method based on the sliding window is used to monitor the vulnerability disclosure frequency in the component warehouse update log in real time, and the abnormal feature dimension is dynamically expanded according to the newly emerged attack vector characteristics to generate a dynamic threshold that matches the current threat environment; The anomaly detection model is combined with dynamic thresholds to analyze the risk score mutation phenomenon during component version iteration in the dependency chain, and to identify compound threat points that simultaneously meet the requirements of topological structure anomaly, excessive vulnerability density, and delayed community repair response. When the vulnerability density of a component dependency path exceeds three times the standard deviation of similar components in the same period, it is marked as a supply chain threat point and a supply chain threat point list is generated.

5. The method according to claim 1, characterized in that According to the risk report, a comprehensive risk map is drawn using a multi-dimensional visualization algorithm to obtain a risk view, including: According to the risk report, integrating dependency chain topology data, threat point spatial location information and risk score time series, constructing a structured dataset containing component coordinates, risk level labels and threat propagation direction vectors; Using a multi-dimensional visualization algorithm, the dependency chain is mapped as a connection line in three-dimensional space, the risk score is converted into a node size parameter, and the threat point location is marked as a pulse signal marker to generate an interactive spatial distribution map; Based on the visualization, density-aware clustering technology is used to automatically identify high-risk component clustering areas, and risk pattern groups are generated according to the component development team, license type, and vulnerability history similarity, to achieve a hierarchical display of threat propagation paths. Combined with the network topology, a ring layout algorithm is designed to optimize node arrangement, gradient color coding is used to represent the risk trend changes in different time windows, and focus + context visualization technology is integrated to achieve scalable display of large-scale dependent networks, generating a comprehensive risk map with multi-dimensional characteristics of time and space; The intensity of vulnerability transmission between components is displayed through the overlay of thermal maps, and the risk diffusion path in the supply chain is presented using dynamic streamline diagrams. The risk node penetration query function is provided to reveal the underlying threat evidence chain and form an actionable risk view.

6. The method according to claim 1, characterized in that The step of locating the node paths associated between the software components according to the attribute characteristics of the software components to generate a dependency chain in the node path, assigning a weight value to the dependency chain, and generating a weighted dependency chain list includes: Using the parsed software component identification information submitted by users, structural attribute features including version number, build tool type and dependency declaration file format are extracted; According to the attribute characteristics of the software components, a multi-level index model of component dependencies is established, and the explicitly declared dependencies and implicitly derived dependencies are identified through a dependency resolution engine to generate a dependency network including version constraints; Based on the dependency network, the dynamically loaded third-party library call records in the build process are traced, the interface call paths covered by the test cases are analyzed, and the actual effective dependencies are verified in combination with the continuous integration logs to generate dependency chain records accurate to the specific call links; Classify the dependency chain records, divide direct and indirect dependencies according to whether the dependency introduction method is required for the development environment and whether it is mandatory to load at runtime, assign exponentially increasing propagation weights to deeply nested cross-ecological dependencies, and generate a preliminary weighted dependency chain; Based on the risk report, a weight correction rule engine is established to dynamically adjust the weight coefficient according to the reputation rating of the component maintainer, the binary file signature verification status and the security factors of the dependency locked file integrity. The rationality of the weight distribution is verified through dependency propagation simulation, and finally a weighted dependency chain list is generated after multiple rounds of iterative optimization.

7. A multi-source software supply chain intelligent analysis system, characterized in that: include: A parsing module, used to parse the identification information of the software component submitted by the user, and extract the attribute characteristics of the software component from the identification information; A positioning module, used to locate the node paths associated between the software components according to the attribute characteristics of the software components, so as to generate a dependency chain in the node path, assign a weight value to the dependency chain, and generate a weighted dependency chain list, wherein the dependency chain includes a direct dependency chain and an indirect dependency chain; A calculation module, for calculating the risk score of the dependency chain in the weighted dependency chain list by using a Bayesian network model combined with a predefined risk indicator system, and identifying potential supply chain threat points by using an anomaly detection algorithm to generate a risk report including an overall risk score and the supply chain threat points, wherein the supply chain threat points are key links or components that are exploited by attackers to undermine the integrity and security of software during software development, distribution and deployment; A drawing module is used to draw a comprehensive risk map based on the risk report using a multi-dimensional visualization algorithm to obtain a risk view, and integrate a genetic algorithm and a simulated annealing algorithm to explore adjustment suggestions for high-risk areas of the risk view to generate a risk mitigation strategy; wherein the comprehensive risk map is grouped and displayed using a hierarchical clustering algorithm to display components with similar risk patterns, wherein the risk pattern is obtained by similarity judgment of the dependency chain, the supply chain threat point and the risk score, and the risk mitigation strategy is used to reduce the risk score of the software component; The Bayesian network model is used in combination with a predefined risk indicator system to calculate the risk score of the dependency chain in the weighted dependency chain list, including: According to the Bayesian network model, the node paths in the weighted dependency chain list are vectorized using the Node2Vec or GraphSAGE graph embedding algorithm, and the component nodes, edge attributes and topological structures in the dependency chain are mapped to a low-dimensional vector space to obtain a dependency vector; Based on the dependency vector, the number of vulnerabilities, update frequency, and community activity indicators are normalized through a dynamic weight allocation strategy, where the weight of the number of vulnerabilities is dynamically adjusted according to the severity level of the CVE database, the update frequency weight is calculated by exponential decay in combination with the component version iteration cycle, and the community activity weight is generated based on linear interpolation of code submission frequency and issue response time to form a multi-dimensional weighted initial risk score; Using the dependency chain with assigned multi-dimensional weighted initial risk scores, the parent node risk probability in the Bayesian network is used as the feature input of the gradient boosting decision tree. Through feature engineering, component dependency depth, cross-ecological call relationship and license conflict type are extracted as additional features. The greedy algorithm is used to optimize the conditional probability convergence direction when the decision tree splits the node, and the optimized conditional probability is obtained. According to the optimized conditional probability, the Stacking ensemble learning method is used to perform meta-model fusion on the component vulnerability prediction results of the random forest classifier and the risk propagation coefficient prediction results of the LightGBM regressor. The weights of each base model are determined through cross-validation, and the risk assessment score of the dependency chain is output.

8. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-source software supply chain intelligent analysis method as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a multi-source software supply chain intelligent analysis method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intelligent perception finished product repair decision support system

    CN117829554A

  • Software supply chain security assessment method and system based on static analysis

    CN118364462A