Software supply chain situation awareness method and device, electronic equipment and storage medium
By constructing a risk evolution link model and an LSTM network, the problem of proactive defense and situational awareness of dynamic risks in the software supply chain is solved, enabling accurate prediction and proactive defense of software supply chain risks, and improving the accuracy and timeliness of risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA SOFTWARE TESTING CENT
- Filing Date
- 2026-03-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are ill-equipped to cope with the complex and dynamic risk environment in the software supply chain, making it difficult to achieve proactive defense and situational awareness, leading to serious consequences such as data breaches and business interruptions.
By extracting risk characteristics from multi-source heterogeneous data, a risk evolution link model of 'direct risk factors - intermediate events - top events' is constructed. Probabilistic calculation methods are used in conjunction with long short-term memory neural networks (LSTM) for risk prediction and situational awareness.
It enables precise and proactive defense and situational awareness of software supply chain risks, improves the accuracy and timeliness of risk assessment, and can proactively warn of future risk trends, providing forward-looking guidance.
Smart Images

Figure CN122451900A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cybersecurity technology, and in particular to a software supply chain situation awareness method, apparatus, electronic device, and storage medium. Background Technology
[0002] In today's digital age, the software industry is booming, and the software supply chain is becoming increasingly complex, making security a critical challenge. Current security measures in the software supply chain primarily rely on signature matching of known threats. When the software supply chain is attacked, it often takes a considerable amount of time to detect and respond. This can lead to consequences such as data breaches and business disruptions. Therefore, while existing security protection methods can identify known and explicit security vulnerabilities, they are insufficient to cope with increasingly complex dynamic risk environments and cannot achieve true proactive defense and situational awareness. Summary of the Invention
[0003] The purpose of this invention is to provide at least one software supply chain situation awareness method, device, electronic device, and storage medium, which can at least solve the technical problem that existing methods are unable to cope with increasingly complex dynamic risk environments and are unable to achieve true proactive defense and situation awareness, and can at least achieve more accurate and proactive risk defense and situation awareness.
[0004] To address the aforementioned technical problems, at least one embodiment of this application provides a software supply chain situational awareness method, comprising: Extract risk characteristics from multi-source heterogeneous data in the software supply chain to be evaluated; Based on the risk characteristics, the direct risk factors existing in the software supply chain to be evaluated are determined. The probability that the top event is a supply chain risk event is calculated according to the risk evolution link of "direct risk factor - intermediate event - top event". Among them, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. The likelihood of future risk events in the software supply chain under evaluation is perceived based on the periodic changes in the probability of such events.
[0005] At least one embodiment of this application also provides a software supply chain situational awareness device, comprising: The risk feature extraction module is used to extract risk features from multi-source heterogeneous data in the software supply chain to be evaluated; The risk probability calculation module is used to determine the direct risk factors existing in the software supply chain to be evaluated based on the risk characteristics, and calculate the probability that the top event is a supply chain risk event according to the risk evolution link of "direct risk factor - intermediate event - top event". Among them, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. The risk probability prediction module is used to perceive the future risk event situation of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk events.
[0006] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described software supply chain situational awareness method.
[0007] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described software supply chain situational awareness method.
[0008] The software supply chain situational awareness method, apparatus, electronic device, and storage medium provided in this application firstly achieve efficient integration and structured processing of complex information in the software supply chain by extracting risk characteristics from multi-source heterogeneous data, laying the foundation for comprehensive risk identification. Secondly, a risk evolution link model of "direct risk factors - intermediate events - top events" is constructed, and a probabilistic calculation method is used to quantify the risk transmission process. This design not only clearly reveals the cascading amplification mechanism of risks in the supply chain, but also accurately assesses the contribution of each link to the overall risk, upgrading risk assessment from qualitative judgment to quantitative analysis. Finally, by analyzing the periodic changes in the probability of supply chain risk events and making future predictions, a leap from static assessment to dynamic awareness is achieved. The method in this embodiment can proactively warn of future risk trends based on historical situation, providing forward-looking guidance for security decisions, effectively making up for the shortcomings of traditional methods in risk prediction and proactive defense, and improving the accuracy, timeliness, and proactivity of situational awareness.
[0009] In some optional embodiments, calculating the probability that the top event is a supply chain risk event according to the risk evolution chain of "direct risk factor - intermediate event - top event" includes: The second probability of occurrence of the corresponding intermediate event is calculated based on the first probability of occurrence of each of the direct risk factors and the first weight corresponding to each of the direct risk factors. The probability that the top event is a supply chain risk event is calculated based on the second occurrence probability of each intermediate event and the second weight corresponding to each intermediate event.
[0010] In this embodiment, the complex risk evolution process is decomposed into a hierarchical structure of "direct risk factors - intermediate events - top events," making the risk assessment process transparent. Weight parameters can be dynamically adjusted based on the characteristics of different software supply chains, historical data, or expert experience, enabling the model to adapt to diverse assessment scenarios. When the importance of certain risk factors changes over time, only the corresponding weights need to be adjusted to update the model, without needing to reconstruct the entire assessment system.
[0011] In some optional embodiments, the method further includes: Based on the changing trend of the probability of future risk events in the software supply chain to be evaluated, adjust the first probability of occurrence of the corresponding direct risk factor; The probability that the current top event is a supply chain risk event is recalculated based on the adjusted first probability of occurrence.
[0012] In this embodiment, the system no longer relies on static, prior risk probabilities, but instead proactively backtracks and adjusts the current assessment values of underlying risk factors based on predictions of future trends. This enables the risk assessment model to have self-learning and real-time evolution capabilities, allowing it to more sensitively reflect the dynamic changes in the supply chain environment and improve the timeliness and accuracy of the assessment.
[0013] In some optional embodiments, the method further includes: The first weight and / or the second weight are adjusted based on the difference between the predicted probability and the actual probability of future risk events in the software supply chain to be evaluated.
[0014] In this embodiment, the risk patterns of the software supply chain evolve over time. When systematic biases occur in the forecast, weight adjustments can promptly correct the model's misjudgment of the importance of certain risk factors or intermediate links, ensuring that the model can adapt to new threat situations and maintain long-term forecast accuracy.
[0015] In some optional embodiments, an initial value for the first occurrence probability of each of the direct risk factors is determined based on the ratio of historical occurrences to the total number of samples.
[0016] In this embodiment, historical experience data is digitized, and objective data is used to replace subjective guesses, thereby laying a stable, reliable and interpretable initial foundation for the entire dynamic risk assessment system.
[0017] In some optional embodiments, the probability that the top event is a supply chain risk event is calculated according to the following formula:
[0018] Specifically, when calculating the second probability of occurrence of the intermediate event, The second probability of occurrence, The first probability of occurrence of direct risk factors. As the first weighting of direct risk factors, The number and types of direct risk factors that led to the intermediate event; When calculating the probability that the top event is a supply chain risk event, Let be the probability that the top event is a supply chain risk event. The second probability of occurrence of the intermediate event. The second weight for intermediate events, The number of types of intermediate events that affect the top event.
[0019] In this embodiment, the same calculation formula is recursively applied to the two layers of "direct risk factors - intermediate events" and "intermediate events - top events". This uniformity makes the modeling structure clear, highly scalable, and better able to adapt to more complex supply chain scenarios.
[0020] In some optional embodiments, the step of sensing the future risk event trend of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk events includes: Obtain a multidimensional time series dataset within a historical window of a preset length. The multidimensional time series dataset includes a probability sequence of the top event being a supply chain risk event, a probability sequence of direct risk factors, and a probability sequence of intermediate events. The multidimensional time series dataset is input into a preset long short-term memory neural network to obtain the predicted probability that the top event in the future is a supply chain risk event, as well as the predicted probabilities of the intermediate events and the direct risk factors corresponding to the top event.
[0021] In this embodiment, the LSTM (Long Short-Term Memory Neural Network) not only outputs the predicted probability of the future top event, but also simultaneously outputs the predicted probabilities of intermediate events and direct risk factors. This allows the system to predict which underlying factors and intermediate links will be the first to show anomalies before the risk of the top event actually increases, thus intervening at the source or early stage of the risk chain, rather than passively waiting for the final event to occur. This gives the entire situational awareness system a more refined self-optimization capability. Attached Figure Description
[0022] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0023] Figure 1 This is a flowchart of a software supply chain situational awareness method provided in one embodiment of this application. Figure 1 ; Figure 2 This is a flowchart of a software supply chain situational awareness method provided in one embodiment of this application. Figure 2 ; Figure 3 This is a schematic diagram of a software supply chain situational awareness device provided in another embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.
[0025] To facilitate understanding of the embodiments of this application, relevant content regarding security risks in the software supply chain will be introduced first.
[0026] As the software supply chain becomes increasingly complex, existing security measures have revealed many pain points: The difficulty in tracking risks associated with open-source components is a major challenge. Open-source software is widely used in the software supply chain, but its diverse origins and numerous versions make it difficult for developers to fully understand its security status. Attackers may exploit known vulnerabilities in open-source components or inject malicious code into them, and developers often struggle to detect and track these risks in a timely manner. The security capabilities of third-party vendors vary widely. Software supply chains often rely on multiple third-party vendors for components and services; however, these vendors differ significantly in their security measures and management levels. Some vendors may lack robust security mechanisms, making them vulnerable to attackers and potentially leading to supply chain disruptions or software tampering.
[0027] Delayed threat response is also a weakness of existing protections. Traditional security measures rely primarily on signature matching of known threats, which is insufficient for responding to new attacks and zero-day vulnerabilities. When the software supply chain is attacked, it often takes a long time to detect and respond, which can lead to serious consequences such as data breaches and business disruptions.
[0028] Therefore, while existing security protection methods can identify known and explicit security vulnerabilities, they are insufficient to cope with increasingly complex dynamic risk environments and cannot achieve true proactive defense and situational awareness.
[0029] To address the technical problems of existing methods being unable to cope with increasingly complex dynamic risk environments and failing to achieve true proactive defense and situational awareness, this invention proposes a software supply chain situational awareness method. The implementation details of the software supply chain situational awareness method in this embodiment are described below. The following content is only for ease of understanding and is not necessary for implementing this solution.
[0030] Example 1: The software supply chain situation awareness method of this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. Its specific process can be as follows: Figure 1 As shown, it includes: Step 110: Extract risk characteristics of multi-source heterogeneous data in the software supply chain to be evaluated; Specifically, the data comes from multiple sources and is heterogeneous, including software code, component dependency data, development log data, vendor data, runtime behavior data, and compliance documents. Different collection methods are used for different types of data during the data acquisition process. For example, component dependency data can be obtained by parsing metadata such as the POM file of the Maven repository and the package.json file of npm. Vendor data access must meet compliance requirements, and structured documents such as SLA performance records and security audit reports are obtained through standardized RESTful interfaces. Meanwhile, unstructured vendor evaluation reports are converted into searchable text using OCR technology.
[0031] When extracting risk features, three main steps can be used: basic feature engineering, advanced deep learning extraction, and feature optimization. These steps can transform multi-source heterogeneous raw data (such as code snippets, log text, component metadata, etc.) into structured, low-noise, and highly predictive risk features, providing key inputs for subsequent risk assessment and prediction.
[0032] The basic feature engineering can include static feature construction, dynamic feature capture, and feature preprocessing. Static features are static, non-time-series attributes of entities such as components, code, and suppliers in the software supply chain, extracting features that reflect their "inherent risk potential" without relying on runtime data. Dynamic features are dynamic behaviors (such as CI / CD processes and network interactions) and time-series changes in the software supply chain, extracting features that reflect "dynamic risk trends" while preserving time-series characteristics. Feature preprocessing addresses issues such as inconsistent dimensions, redundant categories, and high noise in the original features, transforming them into standardized, low-noise features to ensure the model can be used directly.
[0033] Advanced deep learning extraction refers to the extraction of hidden, nonlinear, and higher-order risk features from high-dimensional, unstructured data (such as 10TB-level CI / CD logs and multimodal attack data) where basic feature engineering struggles to capture complex patterns. For example, LSTM networks, autoencoders, and attention mechanisms can be used to extract the original higher-order risk features.
[0034] Feature optimization addresses the potential redundancy, low importance, and insufficient timeliness of extracted original features. It requires optimization and selection to retain core, high-value features—those directly involved in risk assessment. Machine learning models and interpretive tools can be used to identify the features that contribute most to risk prediction, or statistical methods and algorithms can be used to eliminate redundant features. Furthermore, risk features need to be dynamically updated in response to changes in attack patterns and business scenarios to avoid "feature drift" that could cause model failure.
[0035] Step 120: Based on the risk characteristics, determine the direct risk factors existing in the software supply chain to be evaluated, and calculate the probability that the top event is a supply chain risk event according to the risk evolution link of "direct risk factor - intermediate event - top event". In this case, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. Specifically, after completing the feature extraction and optimization in step 110, the system obtains a set of high-value, low-noise risk feature vectors. This embodiment first maps these features to specific direct risk factors. Direct risk factors refer to specific indicators that can be independently observed and directly characterize the existence of security vulnerabilities in a certain link of the software supply chain; they are the "initial trigger points" for risk evolution.
[0036] When mapping features to specific direct risk factors, a predefined mapping rule library can be used to map risk features to direct risk factors. For example, risk features such as "component version lag," "vendor security certification level," and "code cyclomatic complexity outliers" can be mapped to direct risk factors such as "open source component vulnerability risk," "vendor qualification risk," and "code quality risk," respectively.
[0037] In addition to direct mapping based on the inherent attributes of risk characteristics, it's also possible to combine industry characteristics, business scenarios, and historical risk event data within the software supply chain. Different industries face varying types and degrees of risk in their software supply chains. For example, the financial industry has extremely high requirements for data security and compliance; for software components involving sensitive information processing, features such as supplier security certification levels and code encryption are given greater weight when mapped to direct risk factors. Conversely, the gaming industry prioritizes software performance and user experience; abnormal code complexity can significantly impact game smoothness, resulting in relatively less weight when mapped to code quality risks. Furthermore, by analyzing historical risk event data, it's crucial to identify which risk characteristics frequently occurred and had a significant impact in past events. Strengthening the mapping relationship between these characteristics and direct risk factors ensures that the mapping rules accurately reflect the actual situation.
[0038] The risk evolution chain of "direct risk factor - intermediate event - top event" essentially constructs a three-layer cascaded Bayesian network topology to simulate the evolution path of risk from micro to macro. The bottom layer is the direct risk factor node layer, and the node set is: { , , ..., }, corresponding to the identified direct risk factors. The intermediate event node layer has the following node set: { , , ..., This represents an intermediate security event type with a clear semantic meaning, triggered by a combination of direct risk factors. For example: intermediate events. Attack path formation = f( A component vulnerability exists. Malicious code injection, (CI / CD pipeline configuration error). Top event node layer, node set: {T}, representing the supply chain risk events of ultimate concern. The probability of occurrence of the top event T is determined by all intermediate events. ... A joint decision.
[0039] After constructing the risk evolution chain, simulation verification is an essential step. By simulating changes in direct risk factors under different scenarios, the occurrence of intermediate and culminating events is observed to verify whether the risk evolution chain accurately reflects the evolution path of risk from micro to macro levels. If the simulation results deviate significantly from reality, the reasons are analyzed, and the risk evolution chain is adjusted. For example, the definition of intermediate events may be adjusted, the connection relationships or weights between nodes may be changed, until the simulation results match reality, ensuring that the risk evolution model can accurately predict the probability of supply chain risk events.
[0040] Step 130: Based on the periodic changes in the probability of the supply chain risk events, perceive the future trend of risk events in the software supply chain to be evaluated.
[0041] Specifically, the cyclical changes in the probability of software supply chain risk events can be influenced by various factors, such as business seasonality, technology update cycles, and changes in industry policies. Therefore, multi-period analysis should be conducted when making predictions. This involves collecting supply chain risk event probability data at different time scales (e.g., daily, weekly, monthly, quarterly, and yearly) and analyzing their cyclical patterns. For example, a surge in demand for certain businesses during specific seasons may increase pressure on the software supply chain, raising the probability of risk events; similarly, during technology upgrades, vulnerabilities in components of older technologies may increase the probability of risk events. Multi-period analysis provides a more comprehensive understanding of the cyclical characteristics of probability changes, improving the accuracy of predictions.
[0042] The system can present forecast results in intuitive visualizations, such as charts and maps, allowing decision-makers to quickly understand the probability distribution and trends of risk events. It also provides decision support tools, such as a risk response suggestion generator and risk warning threshold settings. Based on the forecast results and the decision-maker's risk preferences, it automatically generates risk response suggestions to help them formulate reasonable risk response strategies. For example, when a high probability of high-risk events occurring in a region's software supply chain is predicted, the risk response suggestion generator can automatically generate suggestions such as adding backup suppliers or strengthening security audits, allowing decision-makers to make informed decisions.
[0043] In summary, the software supply chain situational awareness method provided in this embodiment firstly achieves efficient integration and structured processing of complex information in the software supply chain by extracting risk characteristics from multi-source heterogeneous data, laying the foundation for comprehensive risk identification. Secondly, it constructs a risk evolution chain model of "direct risk factors—intermediate events—top events," and uses probabilistic calculation methods to quantify the risk transmission process. This design not only clearly reveals the cascading amplification mechanism of risks in the supply chain but also accurately assesses the contribution of each link to the overall risk, upgrading risk assessment from qualitative judgment to quantitative analysis. Finally, by analyzing the periodic changes in the probability of supply chain risk events and making future predictions, it achieves a leap from static assessment to dynamic awareness. The method in this embodiment can proactively warn of future risk trends based on historical situational awareness, providing forward-looking guidance for security decisions, effectively making up for the shortcomings of traditional methods in risk prediction and proactive defense, and improving the accuracy, timeliness, and proactivity of situational awareness.
[0044] In some optional embodiments, the step of calculating the probability that the top event is a supply chain risk event according to the risk evolution chain of "direct risk factor - intermediate event - top event" includes: calculating the second probability of occurrence of the corresponding intermediate event based on the first probability of occurrence of each direct risk factor and the first weight corresponding to each direct risk factor; and calculating the probability that the top event is a supply chain risk event based on the second probability of occurrence of each intermediate event and the second weight corresponding to each intermediate event.
[0045] Specifically, the probability of occurrence of each direct risk factor can be obtained through various data sources and analysis methods. For example, for the direct risk factor of "the existence of component vulnerabilities," the frequency of occurrence of this risk factor within a certain period can be estimated by regularly scanning open-source component repositories, referring to security vulnerability databases, and analyzing historical vulnerability data.
[0046] In some optional embodiments, the first probability of occurrence of each of the direct risk factors is determined based on the ratio of historical occurrences to the total number of samples. Initial value:
[0047] The first weight reflects the degree of influence of each direct risk factor on a specific intermediate event. It can be determined using expert evaluation, the Analytic Hierarchy Process (AHP), or statistical analysis based on historical data. Taking expert evaluation as an example, experts in the field of software supply chain security are invited to score the importance of each direct risk factor in causing the intermediate event based on their experience and knowledge. The scoring results are then normalized to obtain the first weight of each direct risk factor. For emerging risk characteristics (such as logical defects introduced by AI-generated code), the prior probability is determined using the Delphi method (three rounds of expert scoring).
[0048] Based on the first probability of occurrence of each direct risk factor and its corresponding first weight, the second probability of occurrence of the intermediate event is calculated, and then the second weight of the intermediate event is determined. The second weight reflects the degree of influence of each intermediate event on the top event (supply chain risk event). It can also be determined using methods such as expert evaluation and analytic hierarchy process (AHP). For example, for intermediate events such as "attack path formation" and "data breach triggering," their importance and correlation in leading to the occurrence of the supply chain risk event are analyzed to determine their respective second weights.
[0049] Finally, based on the second occurrence probability of each intermediate event and the second weight corresponding to each intermediate event, the probability that the top event is a supply chain risk event is calculated. Based on the calculated probability that the top event is a supply chain risk event, a detailed risk assessment report is generated. The report includes the probability value of the risk event, the contribution level of each direct risk factor and intermediate event, the risk level classification (e.g., high, medium, low risk), and corresponding risk response recommendations.
[0050] In this embodiment, the complex risk evolution process is decomposed into a hierarchical structure of "direct risk factors - intermediate events - top events," making the risk assessment process transparent. Weight parameters can be dynamically adjusted based on the characteristics of different software supply chains, historical data, or expert experience, enabling the model to adapt to diverse assessment scenarios. When the importance of certain risk factors changes over time, only the corresponding weights need to be adjusted to update the model, without needing to reconstruct the entire assessment system.
[0051] In some optional embodiments, the method further includes: adjusting the first probability of occurrence of the corresponding direct risk factor based on the changing trend of the probability of future risk events in the software supply chain to be evaluated; and recalculating the probability that the current top event is a supply chain risk event based on the adjusted first probability of occurrence.
[0052] Specifically, comparing the actual occurrence of a risk event with the top event probability calculated based on the initial first occurrence probability may reveal discrepancies. This discrepancy could be due to inaccurate estimation of the first occurrence probability of direct risk factors. To make the risk assessment model more accurate and reliable, the first occurrence probability needs to be adjusted according to the actual risk situation.
[0053] In some alternative embodiments, trend analysis and forecasting methods can be used to adjust the initial probability of occurrence: Data on the occurrence of direct risk factors and top events in the software supply chain over a past period are collected, and their trends over time are analyzed. For example, data such as the number of open-source component vulnerabilities and the number of supplier delivery delays each month over the past year are statistically analyzed, and trend charts are plotted to observe their patterns. A suitable forecasting model is then selected based on the trend characteristics of the data.
[0054] In some alternative embodiments, different risk scenarios can be constructed based on an analysis of various internal and external factors that the software supply chain may face in the future. For example, optimistic, neutral, and pessimistic scenarios can be constructed, considering factors such as market fluctuations, policy changes, and technological innovations. Under each scenario, the first probability of occurrence of each direct risk factor is assessed. For example, in a pessimistic scenario, due to increased market competition, suppliers may use lower-quality raw materials to reduce costs, thereby increasing the risk of software quality problems and correspondingly increasing the first probability of occurrence of this direct risk factor. A weight is assigned to each scenario based on its likelihood of occurrence, and then the first probability of occurrence of the direct risk factor under each scenario is weighted and averaged to obtain the adjusted first probability of occurrence.
[0055] In this embodiment, the system no longer relies on static, prior risk probabilities, but instead proactively backtracks and adjusts the current assessment values of underlying risk factors based on predictions of future trends. This enables the risk assessment model to have self-learning and real-time evolution capabilities, allowing it to more sensitively reflect the dynamic changes in the supply chain environment and improve the timeliness and accuracy of the assessment.
[0056] In some optional embodiments, the method further includes: adjusting the first weight and / or the second weight based on the difference between the predicted probability and the actual probability of future risk events in the software supply chain to be evaluated.
[0057] Specifically, the software supply chain is affected by a variety of internal and external factors, such as technological innovation, changes in market demand, adjustments to policies and regulations, and changes in the operating conditions of suppliers. Changes in these factors can alter the importance and correlation of different risk factors.
[0058] For example, with the widespread adoption of open-source software, vulnerabilities in open-source components have become a significant risk factor in the software supply chain. If the original weighting settings do not adequately account for this change, it may lead to inaccurate assessments of the risks associated with open-source component vulnerabilities. By adjusting the weights based on the difference between predicted and actual probabilities, the risk assessment model can better adapt to environmental changes and promptly reflect new risk characteristics.
[0059] In some optional embodiments, when adjusting the first weight and / or the second weight, a proportional adjustment method based on the difference can be used: first, calculate the difference between the predicted probability and the actual probability, then divide it by the actual probability to obtain the difference ratio. Generally, the larger the difference ratio, the larger the adjustment magnitude. An adjustment coefficient can be preset, and the difference ratio is multiplied by the adjustment coefficient to obtain the final adjustment magnitude. Finally, the first weight and / or the second weight are adjusted according to the adjustment magnitude. If it is a positive difference (predicted probability is greater than actual probability), the weight of the relevant risk factor or intermediate event needs to be reduced; if it is a negative difference (predicted probability is less than actual probability), the corresponding weight needs to be increased.
[0060] In some optional embodiments, sensitivity analysis can also be used to perform sensitivity analysis on each first and second weight in the risk assessment model. This involves changing the value of each weight and observing the change in the probability of the top event. Sensitivity analysis can determine which weights have a greater impact on the probability of the top event, i.e., those with higher sensitivity.
[0061] In some embodiments, the first weights and / or the second weights can also be adjusted using a pre-trained large model.
[0062] In this embodiment, the risk patterns of the software supply chain evolve over time. When systematic biases occur in the forecast, weight adjustments can promptly correct the model's misjudgment of the importance of certain risk factors or intermediate links, ensuring that the model can adapt to new threat situations and maintain long-term forecast accuracy.
[0063] In some optional embodiments, the probability that the top event is a supply chain risk event is calculated according to the following formula:
[0064] Specifically, when calculating the second probability of occurrence of the intermediate event, The second probability of occurrence, The first probability of occurrence of direct risk factors. As the first weighting of direct risk factors, To determine the number of direct risk factors that led to the intermediate events; when calculating the probability that the top event is a supply chain risk event, Let be the probability that the top event is a supply chain risk event. The second probability of occurrence of the intermediate event. The second weight for intermediate events, The number of types of intermediate events that affect the top event.
[0065] In this embodiment, the same calculation formula is recursively applied to the two layers of "direct risk factors - intermediate events" and "intermediate events - top events". This uniformity makes the modeling structure clear, highly scalable, and better able to adapt to more complex supply chain scenarios.
[0066] In some optional embodiments, perceiving the future risk event trend of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk event includes: acquiring a multi-dimensional time series dataset within a historical window of a preset length, the multi-dimensional time series dataset including a probability sequence of the top event being a supply chain risk event, a probability sequence of direct risk factors, and a probability sequence of intermediate events; inputting the multi-dimensional time series dataset into a preset long short-term memory neural network, and processing it to obtain the predicted probability of the top event being a supply chain risk event within a future period, as well as the predicted probabilities of the intermediate events and the direct risk factors corresponding to the top event.
[0067] Specifically, this embodiment uses the Noisy-OR model to calculate the probabilities of direct risk factors, intermediate events, and top events, and employs a pre-defined Long Short-Term Memory (LSTM) neural network as the supply chain situation awareness layer. The output of the Noisy-OR model directly constitutes the basic data for the supply chain situation awareness layer. 1. Top event probability time series The Noisy-OR model outputs the probability of the top event, "supply chain security incident," calculated at a time granularity (e.g., daily / weekly). , forming a time series { , , ..., For example, by updating the top event probability daily using the Noisy-OR model, a sequence of the past 30 days {0.3, 0.4, ..., 0.68} is obtained. This sequence is directly used as input to the LSTM model in the situational awareness layer to predict the probability trend of the next T steps. , ..., The independence assumption of the Noisy-OR model ensures efficient probability calculation, supports high-frequency (e.g., daily) updates, and meets the data density requirements of time series forecasting.
[0068] 2. Probability distribution characteristics of risk factors The Noisy-OR model takes the following inputs: probabilities are calculated using maximum likelihood estimation, and outputs probability values for several direct risk factors (such as "component vulnerability exists" (P=0.7), "CI / CD configuration error" (P=0.6)), forming a feature vector. .
[0069] Connection logic: With the top event probability sequence { , , ..., The LSTM model is spliced together as a multi-dimensional input feature for the situational awareness layer. For example, the LSTM model simultaneously receives the probability of the top event in the past 7 days and the probability of risk factors in the current day, which improves the sensitivity of the prediction model to risk-driven factors (such as "when the probability of component vulnerability is > 0.5, the probability of the top event in the next 3 days increases by 80%").
[0070] The output of the Noisy-OR model needs to be transformed into an input format that can be directly used by the situation awareness layer through temporal feature engineering. The core steps include: 1. Time window division and sequence construction Fixed-period probability calculation: Run the Noisy-OR model according to business needs (e.g., daily) to output the probabilities of the top event and intermediate events (e.g., the probabilities of "business interruption" and "data leakage"), forming a multi-dimensional time series dataset. .
[0071] Sliding window input: For the prediction time t, construct the input vector using probability data from the past k time windows, such as... This allows us to predict the probability trend over the next k steps. For example, when k=7, we can use the probability of the top event over the past 7 days and the probability of the current risk factor to predict the trend over the next 7 days.
[0072] 2. Temporal correlation of the probability of risk transmission paths The Noisy-OR model outputs the conditional probability of intermediate events, i.e., the prediction of the second weight (such as "attack path formation → system intrusion" (P=0.8)) which can quantify the intensity of risk transmission.
[0073] Weight adjustment: Predict the conditional probability of future intermediate events using LSTM. For example, assign high weight to the time window of "sudden increase in the probability of attack path formation" to strengthen the driving role of this event in future risks.
[0074] Model Collaboration: Dynamic Feedback Between the Prediction Layer and the Noisy-OR Model The output of the situational awareness layer dynamically adjusts the parameters of the Noisy-OR model, forming a closed loop of "evaluation-prediction-reevaluation": 1. Prediction results feed back into prior probabilities If the situational awareness layer (LSTM) predicts that the probability of a zenith event in the next 7 days continues to rise (e.g., from 0.68 to 0.75), then the prior probabilities of key risk features in the Noisy-OR model are dynamically increased (e.g., the prior probability of "released more than 3 years without update" is increased from 0.32 to 0.45), and the current risk is recalculated to ensure that the assessment results are linked to future trends.
[0075] 2. Time-series optimization of risk factor weights Adjusting the impact weights of the Noisy-OR model based on prediction error feedback For example, if the predicted contribution of "Undisclosed Vulnerability Response SLA" is higher than the actual risk contribution, it can be corrected using gradient descent. This improves the temporal stability of the Noisy-OR model output, thereby optimizing the input quality of situational awareness.
[0076] In this embodiment, the LSTM (Long Short-Term Memory Neural Network) not only outputs the predicted probability of the future top event, but also simultaneously outputs the predicted probabilities of intermediate events and direct risk factors. This allows the system to predict which underlying factors and intermediate links will be the first to show anomalies before the risk of the top event actually increases, thus intervening at the source or early stage of the risk chain, rather than passively waiting for the final event to occur. This gives the entire situational awareness system a more refined self-optimization capability.
[0077] Example 2: Based on the above embodiments, this embodiment provides an application example of a software supply chain situational awareness method. A schematic diagram of the overall process of this method is shown below. Figure 2 As shown, the overall architecture of the method comprises four key stages: data acquisition, feature extraction, risk assessment, and situational awareness. First, the data acquisition layer collects heterogeneous data from multiple sources, providing a foundation for subsequent analysis. Next, the feature extraction layer extracts valuable features from the collected data. The risk assessment layer quantifies the risks in the software supply chain based on these features. Finally, the situational awareness layer uses historical data and the current situation to perceive potential future risks. Simultaneously, the situational awareness results can be fed back to the risk assessment, feature extraction, and data acquisition stages, providing collaborative guidance and feedback on data acquisition categories, frequency, methods, feature engineering parameter settings, and risk assessment Bayesian network parameters.
[0078] The specific implementation methods for the feature extraction and risk assessment stages are similar to those in Example 1, and will not be repeated here to avoid duplication.
[0079] Software supply chain risks exhibit multi-dimensional penetration characteristics, and a panoramic identification system can be constructed from four dimensions: "supplier-component-process-third party". The supplier dimension focuses on the entity's qualifications and fulfillment capabilities, including transparency of security practices (e.g., undisclosed vulnerability response SLAs), continuous maintenance capabilities (e.g., "zombie components" with an annual update frequency of <2 times), and geopolitical risks (e.g., suppliers from sanctioned regions). The component dimension covers technical and compliance risks, typically including open-source vulnerabilities (CVSS score ≥9.0 associated with CVE numbers), malicious code injection (e.g., hidden backdoor functions), and license conflicts (e.g., GPL-licensed code used in commercial products). Process risks are rooted in the entire development and operations process; CI / CD pipeline configuration errors (e.g., unverified third-party image pulls), insufficient code review coverage (<30%), and key management oversights (e.g., hard-coded API keys) are all high-risk points. Third-party risks involve ecosystem collaboration, including cloud service provider data breaches, backdoor code introduced through outsourced development, and supply chain finance fraud.
[0080] Different industries exhibit significant differences in their sensitivity to risk factors. The financial sector has zero tolerance for compliance risks (such as PCI DSS violations), while internet companies are more concerned with rapid response to component vulnerabilities; critical infrastructure industries such as energy and transportation need to additionally assess the physical impact of supply chain disruptions (such as industrial software poisoning causing production line shutdowns). By adjusting the weights of risk factors according to industry needs, the accuracy of risk identification is improved, and the false positive rate is reduced.
[0081] Table 1 below shows the risk features extracted in this embodiment.
[0082] Table 1: Risk Characteristics Illustration
[0083] Based on Table 1, following the architecture of "risk characteristics → direct risk factors → intermediate events → top event," 12 direct risk factors are identified, including "component vulnerability existence" and "permission configuration error." The intermediate layer defines 6 intermediate events, such as "attack path formation" and "defense mechanism failure." The top event is "supply chain security incident occurrence." Directed edges between nodes represent causal relationships; for example, the conditional probability of "component vulnerability existence → attack path formation" is set to 0.75 (based on historical attack case statistics). The 12 direct risk factors (belonging to the four dimensions of "supplier-component-process-third party") trigger 6 intermediate results through individual or combined effects. The model uses a directed acyclic graph (DAG) to depict the risk transmission path, divided into "risk characteristics → direct risk factors → intermediate events → top event," forming a complete risk evolution chain.
[0084] In some optional embodiments, intermediate events can be categorized into multiple levels, from low to high. For example, they can be divided into basic intermediate events, intermediate intermediate events, and advanced intermediate events. Basic intermediate events are triggered solely by direct risk factors; intermediate intermediate events are triggered by both basic intermediate events and direct risk factors; and advanced intermediate events can be triggered by both intermediate intermediate events and direct risk factors. Table 2 below shows the correspondence between direct risk factors and intermediate events.
[0085] Table 2: Correspondence between direct risk factors and intermediate events
[0086] For the prior probability, or first probability of occurrence, of the direct risk factor, the probability is calculated using maximum likelihood estimation:
[0087] For emerging risk characteristics (such as logical flaws introduced by AI-generated code), prior probabilities are determined using the Delphi method (3 rounds of expert scoring). The conditional probability table (CPT) is transformed, and a Noisy-OR model is used to simplify the number of parameters. The Noisy-OR model is also used to describe the impact of parent nodes on child nodes. Assuming that parent nodes act independently, the probability that the top event is a supply chain risk event is calculated using the following formula:
[0088] Specifically, when calculating the second probability of occurrence of the intermediate event, The second probability of occurrence, The first probability of occurrence of direct risk factors. As the first weighting of direct risk factors, To determine the number of direct risk factors that led to the intermediate events; when calculating the probability that the top event is a supply chain risk event, Let be the probability that the top event is a supply chain risk event. The second probability of occurrence of the intermediate event. The second weight for intermediate events, The number of types of intermediate events that affect the top event.
[0089] Bayesian networks based on the Noisy-OR model provide core input to the situation awareness layer by outputting time series of risk event probabilities and dynamic probability distributions of key risk factors, achieving a seamless connection between "current risk quantification and future trend prediction." The specific path is as follows: A. Output of the Noisy-OR model: Probability time series and dynamic features The Noisy-OR model in Bayesian networks is used to calculate the probabilities of underlying risk factors, intermediate events, and top events (such as "attack path formation" and "supply chain security incident occurrence"). Its output directly constitutes the basic data for the situational awareness layer. 1. Top event probability time series The Noisy-OR model outputs the probability of the top event, "supply chain security incident," calculated at a time granularity (e.g., daily / weekly). , forming a time series { , , ..., For example, by updating the top event probability daily using the Noisy-OR model, we obtain the sequence {0.3, 0.4, ..., 0.68} for the past 30 days.
[0090] Connection logic: This sequence is directly used as input to the LSTM model in the situational awareness layer to predict the probabilistic trend in the next T steps. , ..., The independence assumption of the Noisy-OR model ensures efficient probability calculation, supports high-frequency (e.g., daily) updates, and meets the data density requirements of time series forecasting.
[0091] 2. Probability distribution characteristics of risk factors The Noisy-OR model takes as input a probability calculated using maximum likelihood estimation and outputs the probability values of 12 direct risk factors (such as "component vulnerability exists" (P=0.7) and "CI / CD configuration error" (P=0.6)), which form the feature vector. .
[0092] Connection logic: With the top event probability sequence { , , ..., The LSTM model is spliced together as a multi-dimensional input feature for the situational awareness layer. For example, the LSTM model simultaneously receives the probability of the top event in the past 7 days and the probability of risk factors in the current day, which improves the sensitivity of the prediction model to risk-driven factors (such as "when the probability of component vulnerability is > 0.5, the probability of the top event in the next 3 days increases by 80%").
[0093] B. Temporal Modeling: Transformation from Noisy-OR Probabilities to Predicted Features The output of the Noisy-OR model needs to be transformed into an input format that can be directly used by the situation awareness layer through temporal feature engineering. The core steps include: 1. Time window division and sequence construction Fixed-period probability calculation: Run the Noisy-OR model according to business needs (e.g., daily) to output the probabilities of the top event and intermediate events (e.g., the probabilities of "business interruption" and "data leakage"), forming a multi-dimensional time series dataset. .
[0094] Sliding window input: For the prediction time t, construct the input vector using probability data from the past k time windows, such as... This allows us to predict the probability trend over the next k steps. For example, when k=7, we can use the probability of the top event over the past 7 days and the probability of the current risk factor to predict the trend over the next 7 days.
[0095] 2. Temporal correlation of the probability of risk transmission paths The Noisy-OR model outputs the prediction of the conditional probability of intermediate events (such as "attack path formation → system intrusion" (P=0.8)) which can quantify the risk transmission strength, i.e., the weights.
[0096] Weight adjustment: Predict the conditional probability of future intermediate events using LSTM. For example, assign high weight to the time window of "sudden increase in the probability of attack path formation" to strengthen the driving role of this event in future risks.
[0097] C. Model Collaboration: Dynamic Feedback Between the Prediction Layer and the Noisy-OR Model The output of the situational awareness layer dynamically adjusts the parameters of the Noisy-OR model, forming a closed loop of "evaluation-prediction-reevaluation": 1. Prediction results feed back into prior probabilities If the situational awareness layer (LSTM) predicts that the probability of a zenith event in the next 7 days continues to rise (e.g., from 0.68 to 0.75), then the prior probabilities of key risk features in the Noisy-OR model are dynamically increased (e.g., the prior probability of "released more than 3 years without update" is increased from 0.32 to 0.45), and the current risk is recalculated to ensure that the assessment results are linked to future trends.
[0098] 2. Time-series optimization of risk factor weights Adjusting the impact weights of the Noisy-OR model based on prediction error feedback For example, if the predicted contribution of "Undisclosed Vulnerability Response SLA" is higher than the actual risk contribution, it can be corrected using gradient descent. This improves the temporal stability of the Noisy-OR model output, thereby optimizing the input quality of situational awareness.
[0099] To verify the effectiveness of the method of the present invention, the following verification indicators and methods were adopted:
[0100] The verification tool uses a real software supply chain attack dataset and simulates third-party component poisoning scenarios. By comparing the verification results with industry standards, it is shown that the method of this invention has significant advantages in risk detection and situational awareness.
[0101] Example 3: Another embodiment of this application relates to a software supply chain situation awareness device. The implementation details of this embodiment are described below. The following details are for ease of understanding and are not essential for implementing this solution. A schematic diagram of the software supply chain situation awareness device in this embodiment can be seen as follows: Figure 3 As shown, it includes a risk feature extraction module 310, a risk probability calculation module 320, and a risk dynamic perception module 330.
[0102] Risk feature extraction module 310 is used to extract risk features from multi-source heterogeneous data in the software supply chain to be evaluated; The risk probability calculation module 320 is used to determine the direct risk factors existing in the software supply chain to be evaluated based on the risk characteristics, and calculate the probability that the top event is a supply chain risk event according to the risk evolution link of "direct risk factor - intermediate event - top event". Among them, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. The risk dynamic perception module 330 is used to perceive the future risk event situation of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk events.
[0103] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.
[0104] In some optional embodiments, the software supply chain situation awareness device can implement the software supply chain situation awareness method described in any of the above embodiments.
[0105] Example 4: Another embodiment of this application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the software supply chain situational awareness method in the above embodiments.
[0106] In this embodiment, the memory and processor are connected via a bus, which can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and the memory together. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be further described in this embodiment. A bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0107] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0108] Example 5: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0109] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.
Claims
1. A software supply chain situational awareness method, characterized in that, include: Extract risk characteristics from multi-source heterogeneous data in the software supply chain to be evaluated; Based on the risk characteristics, the direct risk factors existing in the software supply chain to be evaluated are determined. The probability that the top event is a supply chain risk event is calculated according to the risk evolution link of "direct risk factor - intermediate event - top event". Among them, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. The likelihood of future risk events in the software supply chain under evaluation is perceived based on the periodic changes in the probability of such events.
2. The software supply chain situation awareness method according to claim 1, characterized in that, The calculation of the probability that the top event is a supply chain risk event, based on the risk evolution chain of "direct risk factor - intermediate event - top event", includes: The second probability of occurrence of the corresponding intermediate event is calculated based on the first probability of occurrence of each of the direct risk factors and the first weight corresponding to each of the direct risk factors. The probability that the top event is a supply chain risk event is calculated based on the second occurrence probability of each intermediate event and the second weight corresponding to each intermediate event.
3. The software supply chain situation awareness method according to claim 2, characterized in that, The method further includes: Based on the changing trend of the probability of future risk events in the software supply chain to be evaluated, adjust the first probability of occurrence of the corresponding direct risk factor; The probability that the current top event is a supply chain risk event is recalculated based on the adjusted first probability of occurrence.
4. The software supply chain situation awareness method according to claim 2, characterized in that, The method further includes: The first weight and / or the second weight are adjusted based on the difference between the predicted probability and the actual probability of future risk events in the software supply chain to be evaluated.
5. The software supply chain situation awareness method according to claim 2, characterized in that, The initial value of the first occurrence probability of each of the aforementioned direct risk factors is determined based on the ratio of the historical occurrence frequency to the total sample size.
6. The software supply chain situational awareness method according to claim 2, characterized in that, The probability that the top event is a supply chain risk event is calculated using the following formula: Specifically, when calculating the second probability of occurrence of the intermediate event, The second probability of occurrence, The first probability of occurrence of direct risk factors. As the first weighting of direct risk factors, The number and types of direct risk factors that led to the intermediate event; When calculating the probability that the top event is a supply chain risk event, Let be the probability that the top event is a supply chain risk event. This represents the second probability of occurrence of the intermediate event. The second weight for intermediate events, The number of types of intermediate events that affect the top event.
7. The software supply chain situational awareness method according to any one of claims 1-6, characterized in that, The method of perceiving the future risk event trend of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk events includes: Obtain a multidimensional time series dataset within a historical window of a preset length. The multidimensional time series dataset includes a probability sequence of the top event being a supply chain risk event, a probability sequence of direct risk factors, and a probability sequence of intermediate events. The multidimensional time series dataset is input into a preset long short-term memory neural network to obtain the predicted probability that the top event in the future is a supply chain risk event, as well as the predicted probabilities of the intermediate events and the direct risk factors corresponding to the top event.
8. A software supply chain situational awareness device, characterized in that, include: The risk feature extraction module is used to extract risk features from multi-source heterogeneous data in the software supply chain to be evaluated; The risk probability calculation module is used to determine the direct risk factors existing in the software supply chain to be evaluated based on the risk characteristics, and calculate the probability that the top event is a supply chain risk event according to the risk evolution link of "direct risk factor - intermediate event - top event". Among them, one or more of the direct risk factors lead to a type of intermediate event, and each intermediate event jointly determines the probability that the top event is a supply chain risk event. The risk probability prediction module is used to perceive the future risk event situation of the software supply chain to be evaluated based on the periodic changes in the probability of the supply chain risk events.
9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the software supply chain situational awareness method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the software supply chain situational awareness method as described in any one of claims 1 to 7.