Satellite safety assessment method
By employing preliminary and detailed system safety assessment methods, potential hazards are identified, their severity is classified, and failure mode and effects analysis and fault tree analysis are conducted. This addresses the lack of safety assessment for satellite systems and ensures the safety of satellite systems under various operational scenarios.
Patent Information
- Application Number
- CN202511241667.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-12
AI Technical Summary
The lack of a comprehensive satellite safety assessment methodology makes it impossible to ensure the safety of satellite systems under various operational scenarios.
A satellite safety assessment method is provided, including a preliminary system safety assessment and a detailed system safety assessment. By identifying potential hazards, classifying the severity of hazards, assigning safety requirements, and conducting failure mode and effects analysis and fault tree analysis, the method quantitatively analyzes whether the safety requirements are met.
It enables a comprehensive safety assessment of the satellite system, ensuring the safety of the satellite in various operating scenarios and meeting relevant standards and user requirements.
Smart Images

Figure CN121124908A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of satellite security technology, and in particular to a satellite security assessment method. Background Technology
[0002] Satellite systems play a vital role in many fields of modern society, including communication, navigation, and weather monitoring. The security of satellite systems directly affects the reliability and continuity of these applications. Satellite systems are complex, comprising various payloads responsible for performing specific tasks and various platform systems that support these payloads, ensuring stable satellite operation in the space environment. Currently, there is a lack of comprehensive security assessment methods to ensure the safety of such complex systems under various operational scenarios. Summary of the Invention
[0003] To address at least some of the problems mentioned above in the prior art, the present invention provides a satellite security assessment method, comprising:
[0004] Define the scope and objectives of the assessment;
[0005] Collect information about the satellite system;
[0006] Based on the assessment scope and objectives, a preliminary system security assessment is conducted using information from the satellite system. In the early stages of satellite system design, security requirements are determined and allocated to each subsystem. A preliminary analysis of potential hazards and their impact is also performed.
[0007] Based on the assessment scope and objectives, a detailed system security assessment is conducted using information from the satellite system. During the design refinement phase, the system failure modes and impacts are further analyzed to determine the risk level and failure probability, and a quantitative analysis is performed to determine whether the security requirements are met.
[0008] Furthermore, defining the scope and objectives of the assessment includes:
[0009] The assessment scope covers the entire lifecycle of the satellite, from launch to the end of its lifespan, including the satellite itself and its ground support systems; and
[0010] The assessment objective is to ensure that the satellite's safety risks are acceptable under specified mission scenarios, meeting relevant standards and user requirements; and
[0011] Collecting information about the satellite system includes gathering design documents and operational data, among which:
[0012] The design documents include the overall satellite design scheme, subsystem design drawings, and technical specifications to clarify the satellite system architecture, functions, and performance indicators.
[0013] Operational data includes historical data on satellite operation in orbit, which is used to analyze problems and potential risks that arise during actual operation.
[0014] Furthermore, the preliminary system security assessment includes:
[0015] Identify potential hazards based on satellite mission and operating environment;
[0016] Classify dangers according to their severity; and
[0017] The overall security requirements are broken down and allocated to each subsystem.
[0018] Furthermore, identifying potential hazards based on satellite mission and operating environment includes:
[0019] Perform functional hazard analysis, analyzing the hazards caused by functional failure one by one;
[0020] Identify single points of failure and the hazards they pose.
[0021] Conduct a preliminary hazard analysis and identify top-level hazards based on the mission scenario;
[0022] Perform event tree analysis to deduce the consequences of the initial event;
[0023] Specific space environment analyses include: space debris collision probability model analysis of collision probability, radiation effect analysis, and surface charging and discharging risk analysis caused by geomagnetic storms.
[0024] Furthermore, classifying dangers according to their severity includes:
[0025] Hazards are classified in descending order of severity, including: catastrophic hazard, hazardous hazard, major hazard, minor hazard, and negligible hazard.
[0026] A catastrophic hazard is a danger that could result in death or complete loss of the system.
[0027] Danger is a risk that could result in serious injury to personnel or severe damage to the system.
[0028] Significant hazards are those that could lead to a substantial degrade of mission capability;
[0029] Minor hazards are those that reduce task efficiency but are recoverable; and
[0030] Negligible hazards are those that have no safety impact.
[0031] The overall security requirements are broken down and allocated to each subsystem, including:
[0032] Based on quantitative allocation and qualitative requirements, the system-level safety objectives are decomposed into subsystems, and the safety objectives of each subsystem are determined. Quantitative allocation is based on fault tree analysis to calculate the contribution of each subsystem, and the contribution is multiplied by the system-level safety objectives to determine the safety objectives of each subsystem.
[0033] Furthermore, a detailed system security assessment includes Failure Mode and Effects Analysis (FMEA) and Fault Tree Analysis (FMA), where FMEA includes:
[0034] Failure mode and effect analysis was performed on each subsystem, component and software of the satellite to determine the impact, possible causes and detection methods of each failure mode.
[0035] The Risk Priority Number (RPN) is calculated based on the severity of the impact of the failure and the probability of its occurrence.
[0036] Furthermore, the calculation method for RPN is: RPN = SEV × OCC × DET, where:
[0037] SEV stands for Severity, which is used to assess the degree of harm that the consequences of a failure pose to the safety of a system, task, or personnel. It is usually divided into 1 to 10 levels, with level 10 being the most severe.
[0038] OCC occurrence rate is used to assess the probability of failure causes occurring. It is usually divided into 1-10 levels, with level 10 being the most frequent.
[0039] DET stands for Detectability, which is used to assess the ability of existing control measures to detect the cause or failure mode before failure occurs. It is divided into 1-10 levels, with level 10 being almost undetectable.
[0040] Furthermore, fault tree analysis includes constructing a fault tree and analyzing the fault tree, wherein:
[0041] Constructing a fault tree involves: using a major hazardous event as the top event and combining it with the logical relationships of the satellite system to construct a fault tree;
[0042] Fault tree analysis includes: qualitative analysis to find the minimum cut set and to identify critical failure modes, including:
[0043] Traverse the fault tree using either the down-row or up-row method to find all the basic events that can lead to the top event and obtain the cut set. A cut set is a set of basic events. If all events in the set occur, then the top event must occur.
[0044] Minimize all the cut sets found. One of the cut sets is the minimum cut set if and only if any one of the basic events is removed. The remaining set is no longer a cut set, that is, it can no longer cause the top event to occur on its own.
[0045] Sort the minimum cut sets, including:
[0046] Order sorting: first-order minimum cut set > second-order minimum cut set > third-order minimum cut set > ... The lower the order, the higher the risk;
[0047] Within a first-order minimal cut set, or among minimal cut sets of the same order, the basic events are sorted according to their failure rate, severity, detectability, and maintainability.
[0048] Based on the minimum cut set sorting results, identify key failure modes, improve the design, evaluate the effectiveness of redundancy, and develop maintenance and testing strategies.
[0049] Furthermore, when performing quantitative analysis, input the basic event failure rate data λ, task time t, common cause failure parameter β, and maintenance strategy;
[0050] Quantitative analysis includes:
[0051] Calculate the probability of the basic event, including: usually assuming a constant failure rate, the probability of the basic event failing within the task time t is P(Basic)≈λ*t;
[0052] The probability of the top event is calculated using the minimum cut set MCS, including:
[0053] The probability of the top event occurring is approximately equal to the union of the probabilities of all minimal cut sets, where the probability of the top event occurring is P(Top)≈ΣP(MCSi)-ΣΣP(MCSi∩MCSj)+ΣΣΣP(MCSi∩MCSj∩MCSk)-...;
[0054] When calculating the probability of occurrence of a minimum cut set containing redundant components, the probability of a common cause failure event must be included in the calculation of the corresponding minimum cut set, where the probability of a common cause failure event is calculated based on the common cause failure parameter β.
[0055] External events are treated as independent basic events, and their occurrence probability is included in the calculation.
[0056] Furthermore, it also includes: verifying whether the satellite system design and implementation meet the safety requirements, and confirming whether the satellite system meets the safety requirements in actual use.
[0057] The present invention has at least the following beneficial effects: The satellite safety assessment method of the present invention includes two stages: preliminary system safety assessment and detailed system safety assessment. In the early stage of satellite system design, the safety requirements of the satellite system are determined and allocated to each subsystem of the satellite system. Potential hazards and their impact are preliminarily analyzed. In the design refinement stage, the system failure modes and impacts are further analyzed, the risk level and failure probability are determined, and the safety requirements are quantitatively analyzed. This method can conduct a comprehensive safety assessment of the satellite system to ensure the safety of the satellite system under various operating scenarios. Attached Figure Description
[0058] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the embodiments of the invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.
[0059] Figure 1 A flowchart of a satellite security assessment method according to an embodiment of the present invention is shown.
[0060] Figure 2 A schematic diagram of a fault tree according to an embodiment of the present invention is shown. Detailed Implementation
[0061] It should be noted that the components in the accompanying drawings may be shown exaggerated for illustrative purposes and may not be to scale.
[0062] In this invention, the various embodiments are merely intended to illustrate the solutions of the invention and should not be construed as limiting.
[0063] In this invention, unless otherwise specified, the quantifiers “a” and “one” do not exclude scenarios involving multiple elements.
[0064] It should also be noted that, in the embodiments of the present invention, only a portion of the parts or components may be shown for clarity and simplicity. However, those skilled in the art will understand that, under the teachings of the present invention, the required parts or components can be added as needed for specific scenarios.
[0065] It should also be noted that within the scope of this invention, the terms "same", "equal", and "equal to" do not mean that the two values are absolutely equal, but allow for a certain reasonable error. In other words, the terms also cover "substantially the same", "substantially equal", and "substantially equal to".
[0066] It should also be noted that in the description of this invention, the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not explicitly or implicitly suggest that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0067] Furthermore, the embodiments of the present invention describe the process steps in a specific order. However, this is only for the convenience of distinguishing each step, and is not a limitation on the order of each step. In different embodiments of the present invention, the order of each step can be adjusted according to the process.
[0068] The satellite architecture consists of a payload and a platform system. The payload is responsible for performing specific tasks, such as the transponder of a communication satellite or the imaging equipment of a remote sensing satellite. The payloads of different types of satellites vary greatly. The platform system provides support for the payload and includes a power system (solar panels, batteries, etc.), an attitude control system (gyroscopes, thrusters, etc.), an orbit control system, a thermal control system, etc., to ensure the stable operation of the satellite in the space environment.
[0069] The satellite's operating environment includes space radiation, micrometeoroids and space debris, and extreme temperatures.
[0070] Space radiation, including high-energy protons and electrons, can cause single-particle flipping of satellite electronic components, affecting the normal function of the equipment.
[0071] Micrometeoroids and space debris: Impacts from tiny particles can cause damage to the surface of a satellite, while impacts from larger debris can lead to severe structural damage or even the disintegration of the satellite.
[0072] Extreme temperatures: Temperatures on the sun-facing side can reach over 100°C, while temperatures on the shaded side can drop to below -100°C, posing a huge challenge to the satellite's thermal control system and material properties.
[0073] A satellite security assessment method includes two phases: Preliminary System Security Assessment (PSSA) and Detailed System Security Assessment (DSSA).
[0074] Preliminary System Safety Assessment (PSSA): In the early stages of satellite system design, system safety requirements are determined and allocated to each subsystem, and potential hazards and their impact are preliminarily analyzed.
[0075] Detailed System Safety Assessment (DSSA): During the design refinement phase, a more in-depth analysis of system failure modes and effects is conducted to determine failure probabilities and risk levels, and to verify whether safety requirements are met. DSSA employs Failure Mode and Effects Analysis (FMEA) and Fault Tree Analysis (FTA).
[0076] Failure Mode and Effects Analysis (FMEA) includes:
[0077] Functional-level FMEA: From the perspective of the overall satellite function, analyze the impact of each functional failure on mission and safety. For example, communication function interruption affects ground communication services and may lead to obstruction of important information transmission.
[0078] Hardware-level FMEA: For satellite hardware components, such as circuit boards and sensors, analyze the impact of individual component failures on system functionality and safety. For example, a power module failure may cause power outages in some satellite equipment.
[0079] Software-level FMEA: For satellite control software, analyze the impact of software faults (such as algorithm errors, logic vulnerabilities, etc.) on satellite operation and safety. For example, erroneous instructions from attitude control software may cause the satellite to lose attitude control.
[0080] The Fault Tree Analysis (FTA) process includes:
[0081] Top event determination: The top event is defined as the failure of a satellite mission or a major safety incident, such as satellite loss of contact or explosion.
[0082] Fault Tree Construction: Based on the system structure and fault logic, identify all possible fault combinations that lead to the top event, determine intermediate and bottom events, and construct a fault tree. For example, a satellite power outage may be caused by multiple factors such as solar panel failure, power transmission line failure, and battery failure.
[0083] Qualitative and quantitative analysis: Qualitative analysis identifies the minimum cut set and determines the key failure mode that leads to the top event; quantitative analysis calculates the probability of the top event and assesses the system's security risks.
[0084] The following is combined Figure 1 This document details the specific steps involved in a satellite security assessment method. One such method includes the following steps:
[0085] Step 1: Define the scope and objectives of the assessment. Specifically, the assessment scope covers the entire lifecycle of the satellite, from launch to the end of its lifespan, including the satellite itself and ground support systems (such as ground stations). The assessment objective is to ensure that the satellite's safety risks are acceptable and meet relevant standards and user requirements under the specified mission scenarios.
[0086] Step 2: Gather information about the satellite system. Specifically, this includes collecting design documents and operational data. Design documents include the overall satellite design scheme, subsystem design drawings, technical specifications, etc., to understand the satellite system architecture, functions, and performance indicators. Collecting operational data includes gathering historical on-orbit data, such as fault records and telemetry data, to analyze problems and potential risks encountered during actual operation.
[0087] Step 3: Based on the assessment scope and objectives, conduct a preliminary system safety assessment (PSSA) based on the information of the satellite system. In the early stages of satellite system design, determine the satellite system safety requirements and allocate them to each subsystem of the satellite system, and conduct a preliminary analysis of potential hazards and their impact.
[0088] Preliminary System Security Assessment (PSSA) includes:
[0089] Hazard identification: Based on the satellite mission and operating environment, identify potential hazards, such as space debris impacts causing satellite structural damage or attitude loss leading to communication interruptions.
[0090] Hazard classification: Hazards are classified according to their severity, including catastrophic, hazardous, major, minor, and negligible. For example, a satellite explosion is a catastrophic hazard, while a localized degradation of equipment performance may be a major hazard.
[0091] Safety requirements allocation: The overall safety requirements are broken down and allocated to each subsystem. For example, the probability of power system failure is required to be below a certain value to ensure the overall safety of the satellite.
[0092] The specific hazard identification process is as follows:
[0093] Functional Hazard Analysis (FHA): Analyzes the hazards caused by functional failures one by one. For example, attitude control failure → satellite roll → solar panel breakage.
[0094] Failure Mode and Effects Analysis (FMEA) of Systems and Subsystems: Identifying single points of failure and the hazards they pose. For example, onboard computer memory errors can lead to attitude data loss and loss of attitude control.
[0095] Preliminary Hazard Analysis (PHA): Identifying top-level hazards based on the mission scenario. Examples include structural cracks caused by vibrations during launch and single-event upsets triggered by on-orbit radiation.
[0096] Event Tree Analysis (ETA): Extrapolates the consequences of an initial event, such as space debris impact → fuel leak → explosion.
[0097] Specific space environment analyses include: space debris collision probability model analysis (using tools such as NASA ORDEM), radiation effect analysis (total dose effect, single-event effect), surface charging and discharging (ESD) risk analysis caused by geomagnetic storms, etc.
[0098] Typical examples of satellite hazards: propellant leaks leading to explosions; solar panel deployment failures triggering energy crises; star sensors being interfered with by strong light causing attitude inaccuracies; thermal control malfunctions causing batteries to overheat and fail.
[0099] The specific methods for hazard classification are as follows:
[0100] Hazards are classified in descending order of severity, including: Level 1 Hazard, Level 2 Hazard, Level 3 Hazard, Level 4 Hazard, and Level 5 Hazard. These five levels correspond to catastrophic, hazardous, major, minor, and no safety effect hazard, respectively.
[0101] Dangers that could result in death or complete loss of a system are classified as catastrophic hazards, such as a satellite exploding in orbit and causing a chain reaction of space debris, or a propulsion system failure leading to the crash of a launch vehicle.
[0102] Hazards that cause serious injury to personnel or severe damage to systems are considered dangerous, such as loss of attitude control leading to ground station communication interruption (critical mission failure), battery thermal runaway causing fire to spread to critical payloads, etc.
[0103] A hazard that significantly degrades mission capability is considered a major hazard, such as a 30% decrease in solar panel output power or a reduction in data transmission rate leading to partial data loss.
[0104] Minor hazards that reduce mission efficiency but are recoverable include things like brief attitude jitter caused by redundant gyroscope switching or automatic repair after a single-event flip of onboard memory.
[0105] Hazards with no safety impact are considered to have negligible safety effects, such as deviations in experimental load calibration parameters or drift of non-critical temperature sensors.
[0106] The specific method for allocating security requirements is as follows:
[0107] The allocation of security requirements specifically refers to assigning system-level security objectives (such as "probability of catastrophic failure of the entire satellite < 1e") to specific targets. -7 The task hour is broken down into subsystems, and the safety indicators for each subsystem are determined.
[0108] The allocation logic is as follows:
[0109] 1. Quantitative Allocation: The contribution of each subsystem is calculated based on Fault Tree Analysis (FTA), and the contribution is multiplied by the system-level safety objective to determine the safety objective of each subsystem. For example, if the propulsion system contributes 60% of the catastrophic risk, its failure rate must satisfy: λ_prop < 0.6 × 1e -7
[0110] 2. Qualitative requirements: Design constraints (e.g., "Power bus must be dual redundant and isolated"), process requirements (e.g., "Critical welds must be 100% X-ray inspected").
[0111] Design constraints: These specify what the system / hardware / software "must be made like," which are rigid requirements for the design outcome.
[0112] Process requirements: Specify "how to do it" to ensure that the goal is achieved. This is a hard requirement for the development, manufacturing, and maintenance processes.
[0113] Table 1. Example of allocation for a satellite subsystem
[0114]
[0115]
[0116] Step 4: Based on the assessment scope and objectives, conduct a detailed system safety assessment (DSSA) using information from the satellite system. During the design refinement phase, further analyze the system failure modes and impacts, determine the risk level and failure probability, and quantitatively analyze whether the safety requirements are met.
[0117] Detailed System Safety Assessment (DSSA) includes Failure Mode and Effects Analysis (FMEA) and Fault Tree Analysis (FTA).
[0118] Failure Mode and Effects Analysis (FMEA) includes comprehensive analysis and risk assessment.
[0119] Comprehensive Analysis: A detailed failure mode and impact analysis (FMEA) is conducted on each subsystem, component, and software of the satellite to determine the impact, possible causes, and detection methods for each failure mode. For example, analyzing a gyroscope failure in the attitude control subsystem may lead to attitude measurement errors. The cause could be gyroscope aging or radiation exposure. Detection methods include regular calibration and monitoring with telemetry data.
[0120] Table 2. Comprehensive analysis of the attitude control subsystem.
[0121]
[0122] Table 3. Comprehensive analysis of the power supply subsystem.
[0123]
[0124]
[0125] Table 4. Comprehensive analysis of the communication subsystem.
[0126]
[0127] Table 5. Comprehensive analysis of spaceborne computers.
[0128]
[0129] Table 6. Comprehensive analysis of the thermal control subsystem.
[0130]
[0131] Table 7. Comprehensive analysis of the propulsion subsystem.
[0132]
[0133]
[0134] Risk assessment: Calculate the Risk Priority Number (RPN) based on the severity of the failure's impact and the probability of its occurrence, and focus on failure modes with high RPN values.
[0135] In the SAE ARP4761A standard, the RPN (Risk Priority Number) is one of the core tools in FMEA / FMECA (Failure Mode, Effects, and Criticality Analysis), used to quantify risk and determine the priority of corrective actions. The following is the specific method for calculating the RPN in satellite safety assessment:
[0136] RPN = SEV × OCC × DET
[0137] SEV stands for Severity, which is used to assess the degree of harm that the consequences of a failure pose to the safety of a system, task, or personnel.
[0138] They are usually classified into levels 1-10 (level 10 being the most severe).
[0139] Satellite example:
[0140] SEV=10: Satellite failure, mission failure, casualties (manned mission).
[0141] SEV=7: Loss of critical functions (such as loss of attitude control).
[0142] SEV=4: Performance degradation (e.g., reduced communication speed).
[0143] SEV=1: No effect.
[0144] OCC stands for Occurrence, which is used to assess the probability of a failure cause occurring.
[0145] They are usually divided into 1 to 10 levels (level 10 is the most frequent).
[0146] Based on historical data, reliability models, or expert judgment:
[0147] OCC=10: Failure is almost inevitable (probability > 1 / 2).
[0148] OCC=7: High probability (e.g., 1 / 100).
[0149] OCC=3: Low probability (e.g., 1 / 10,000).
[0150] OCC=1: Extremely rare (probability <1 / 10) ^6 ).
[0151] DET stands for Detection: used to evaluate the ability of existing control measures to detect the cause or failure mode before failure occurs.
[0152] It is divided into 1 to 10 levels (level 10 is almost undetectable).
[0153] Satellite example:
[0154] DET=10: Undetectable (e.g., no telemetry in deep space environment).
[0155] DET=6: Possibly detected through periodic telemetry analysis.
[0156] DET=3: Real-time sensor automatic alarm.
[0157] DET=1: Design inherent error protection (such as hardware redundancy).
[0158] Fault tree analysis (FTA) includes: constructing the fault tree and analyzing the fault tree.
[0159] Constructing a fault tree involves using a major hazardous event as the top event and combining it with the logical relationships of the satellite system to build the fault tree. For example, using satellite communication interruption as the top event, the path leading to communication interruption caused by failures in each link of the communication link (transmitter, antenna, receiver, etc.) is analyzed.
[0160] Fault tree construction method (top-down):
[0161] 1) Place top event: Write the precisely defined top event in the topmost box of the drawing or FTA software.
[0162] 2) Determine the direct cause of the top event: Question: "What combination of direct causes would lead to a satellite communication outage?" Analyze the system architecture and functional logic.
[0163] Logic gate selection: Typically, the first logic gate in the top event is an OR gate. This is because satellite communication interruptions are usually caused by space segment failure, ground segment failure, or propagation link failure. However, this depends on the system architecture (for example, if there are multiple independent ground stations, an AND gate might be required for all ground stations to fail).
[0164] Intermediate event definition: Decomposing the direct cause into more specific and analyzable intermediate events.
[0165] 3) Determine the direct cause of the intermediate event through step-by-step decomposition:
[0166] For each intermediate event, repeat step 2: identify the combination of direct causes that led to the occurrence of this intermediate event.
[0167] Continue to use AND gates, OR gates, or special gates (such as voting gates and prohibition gates) for logical connections.
[0168] The decomposition principles are as follows:
[0169] Sufficiency: All major direct causes of events at this level must be included.
[0170] Necessity: The combination of listed causes is sufficient to cause this level of event to occur (no critical path omitted).
[0171] Clarity: Every event (intermediate or fundamental) should have a clear and unambiguous definition.
[0172] Atomicity: Decomposition until it cannot be further decomposed, or until the “basic events” for which reliability and failure rate data can be assigned.
[0173] 4) Identifying Basic Events: When decomposing to component-level failure modes (such as "transmitter power amplifier failure," "low-noise amplifier open circuit," "attitude control failure causing excessive antenna pointing deviation," "ground station master server downtime," "ionospheric scintillation causing deep fading"), these are the basic events (leaves) of the fault tree. They require:
[0174] Clearly define the type of fault and the conditions under which it occurs (What kind of fault is it? Under what conditions?).
[0175] Failure rate data (λ) is available.
[0176] It is documented in detail in Failure Mode and Effects Analysis (FMEA).
[0177] 5) Identify external events and common cause failures and add them to the fault tree:
[0178] External events: such as "space debris impact causing antenna damage", "solar storm causing communication disruption", "ground station encountering natural disasters". These are usually added as independent base events or intermediate events.
[0179] Common Cause Failure (CCF): This is particularly critical for redundant designs. For example, the simultaneous failure of power modules in the same batch due to design flaws could lead to the simultaneous failure of redundant transmitters. It requires the use of a CCF model (such as a Beta Factor model) or a specific CCF-based event to represent this. ARP4761A emphasizes the importance of identifying and modeling CCFs.
[0180] 6) Use transfer symbols: For large and complex systems, use transfer-in and transfer-out symbols to connect subtrees and keep the main tree clear.
[0181] 7) Verification and Review:
[0182] Consistency check: Determine whether the logic gates accurately reflect the system design and are consistent with the functional block diagram and FMEA.
[0183] Integrity check: Determines whether all known critical failure paths are included.
[0184] Clarity check: Check whether all event definitions are clear and unambiguous, and whether the logical relationships are easy to understand.
[0185] Independent review: cross-review by different engineers.
[0186] like Figure 2 As shown, the top event in a fault tree related to satellite communication outages is satellite communication outage itself. Satellite communication outages are typically caused by space segment failure, ground segment failure, or propagation link failure. Space segment failures are typically caused by transmitter failure, receiver failure, or antenna failure. Ground segment failures are typically caused by complete failure of the master control station or the backup station. Complete failure of the master control station is caused by server downtime or network failure. Complete failure of the backup station is caused by server downtime or network failure. Propagation link failures are typically caused by ionospheric interference / scintillation or atmospheric attenuation exceeding a threshold.
[0187] Analyze the fault tree: qualitative analysis identifies the minimum cut set and determines the critical failure mode; quantitative analysis calculates the probability of the top event and compares it with the safety target to determine whether the requirements are met.
[0188] The goal of qualitative analysis is to identify all minimal failure combinations (minimum cut sets) that lead to the occurrence of the top event, and to determine the weak points and critical failure modes of the system.
[0189] Qualitative analysis methods include:
[0190] 1) Cut Sets Solution: Traverse the fault tree using either a downward or upward approach to find all combinations (sets) of basic events that can lead to the top event. A cut set is a set of basic events; if all events in the set occur, the top event must have occurred.
[0191] 2) Minimal Cut Sets (MCSs): Minimize all the cut sets found. A cut set is a minimal cut set if and only if removing any one of its basic events results in the remaining set no longer being a cut set (i.e., no longer capable of independently causing the top event). Minimal cut sets represent the weakest failure path of the system.
[0192] 3) Minimum cut set sorting:
[0193] Order ranking: first-order minimal cut set (single point of failure) > second-order minimal cut set (two events fail simultaneously) > third-order minimal cut set > ... The lower the order, the higher the risk is usually.
[0194] Sorting by importance: Within a first-order minimal cut set, or among minimal cut sets of the same order, they can be sorted according to factors such as the failure rate, severity, detectability, and maintainability of the basic events.
[0195] The output and function of the minimum cut set sorting result:
[0196] Critical Failure Mode Identification: The basic events corresponding to all first-order minimal cut sets are single points of failure and represent the highest risk points in the system. They must be eliminated through key design (by increasing redundancy through design changes) or strictly controlled (by improving reliability, increasing testing, and strengthening maintenance).
[0197] Redundancy effectiveness assessment: Analyze the order of the minimum cut sets. If the redundancy design is effective, the relevant failure paths should exhibit high-order minimum cut sets (e.g., second-order or higher). If first-order minimum cut sets exist (e.g., common-cause failure (CCF) events), the redundancy effect is significantly reduced.
[0198] Design improvement basis: After identifying high-risk minimal cut sets (especially low-order minimal cut sets), the design can be improved in a targeted manner (such as increasing redundancy, selecting more reliable devices, improving fault tolerance mechanisms, and increasing health monitoring).
[0199] Maintenance and testing strategy development: Develop more rigorous testing, inspection, and maintenance plans for critical foundational events (especially first-order minimal cut set events).
[0200] Common-cause failure exposure: CCF events are usually low-order minimal cut sets (often first-order), and qualitative analysis can clearly expose their significant threat to system security.
[0201] The goal of quantitative analysis is to calculate the probability of the top event occurring (P(Top)) and compare it with the system's safety objectives (such as the maximum allowable failure probability P(Target)) to assess whether the safety requirements are met.
[0202] When performing quantitative analysis, input the basic event failure rate data (λ), task time (t), common cause failure parameters (β), and maintenance strategy.
[0203] Basic event failure rate data (λ): derived from manuals (e.g., MIL-HDBK-217F, Siemens SN 29500), field data, similar product data, accelerated testing, expert judgment, etc. Data must consider mission profiles (environmental stress, operating time).
[0204] Mission time (t): The time period during which risk needs to be calculated (e.g., duration of a single mission, 1 flight hour).
[0205] Common cause failure parameter (β): If a CCF model such as Beta Factor is used, a corresponding β factor is required (usually derived from industry data or specific analysis).
[0206] Maintenance strategy: Determine whether the system is repairable during the mission to select the appropriate model. (For unrepairable systems, reliability block diagrams / RBD models are commonly used to calculate probabilities; for repairable systems, Markov models or simulations are commonly used to calculate availability / unavailability).
[0207] Quantitative analysis methods (for unrepairable systems, calculating the probability of failure during the mission):
[0208] 1) Calculate the probability of basic events: Usually, a constant failure rate is assumed. The probability of a basic event failing within the task time t is P(Basic)≈λ*t (approximately effective when λt<<1).
[0209] 2) Calculate the probability of the top event using the minimum cut set MCS:
[0210] The probability of the top event is approximately equal to the union of the probabilities of all minimal cut sets.
[0211] Since minimal cut sets are usually not mutually exclusive (they have intersections), direct addition will overestimate their sums. An accurate or approximate calculation is required using the inclusion-exclusion principle or a more efficient disjoint sum algorithm (such as SDP - Sum of Disjoint Products). The FTA software handles this process automatically.
[0212] P(Top)≈ΣP(MCSi)-ΣΣP(MCSi∩MCSj)+ΣΣΣP(MCSi∩MCSj∩MCSk)-...
[0213] For high-order low-probability events, the Rare Event Approximation is often used: P(Top) ≈ ΣP(MCSi). This provides a conservative estimate (too large).
[0214] 3) Consider common cause failure: When calculating the probability of occurrence of a minimum cut set containing redundant components, the probability of the common cause failure event (calculated based on the β factor) should be added to the calculation of the corresponding minimum cut set.
[0215] 4) Consider external events: Treat external events as independent basic events, and include their occurrence probability (if quantifiable) in the calculation.
[0216] Outputting and determining the probability of the top event occurring:
[0217] The calculated P(Top) is the probability of the top event (satellite communication interruption) occurring within a specified task time t.
[0218] Compare with security objectives: Compare P(Top) with the security objectives of the system or function (e.g., P(Target) = 1E-5 / Flight Hour).
[0219] If P(Top) <= P(Target): the security requirements are met.
[0220] If P(Top) > P(Target): The security requirement is not met. When the security requirement is not met, the following measures must be taken:
[0221] Redesign to eliminate high-risk MCS (especially first-order MCS); reduce the failure rate of critical underlying events; optimize CCF protection design (physical isolation, design diversity); add additional security measures or degradation modes.
[0222] Importance metric (optional but recommended):
[0223] Fussell-Vesely Importance (FV): The percentage contribution of a given fundamental event or cutset to the probability of the top event. Used to identify elements that have the greatest impact on security.
[0224] Birnbaum Importance: The sensitivity of the top event probability to changes in the probability of a base event. Used to identify components that require the most reliable control.
[0225] Risk Achievement Worth (RAW): Assuming a base event is certain to occur (probability = 1), the probability of the top event increases by a factor of 1. Used to identify high-risk single points of failure.
[0226] Risk Reduction Worth (RRW): The factor by which the probability of the top event decreases if a base event is assumed to never occur (probability = 0). Used to identify components that can yield the greatest benefit through improvement.
[0227] ARP4761A requires that quantitative analysis is the core means of probabilistic safety objectives (such as the probability of failure requirement of DALA functions). Calculation results and their assumptions (data sources, model selection) must be clearly documented.
[0228] Step 5: Verify whether the satellite system design and implementation meet the safety requirements, and confirm whether the satellite system meets the safety requirements in actual use.
[0229] Through analysis, testing, and simulation, the design and implementation of the satellite system are verified to ensure they meet safety requirements. For example, the radiation resistance of electronic components is tested by simulating the space radiation environment on the ground to verify whether the design specifications are met.
[0230] By using actual operational data and user feedback, we confirm whether the safety of the satellite system meets the requirements in actual use. For example, we collect fault data after the satellite has been in orbit for a certain period of time to assess whether the safety risks are within an acceptable range.
[0231] Verification methods for ensuring that the design and implementation of a satellite system meet safety requirements include:
[0232] Dynamic / in-situ testing: This involves not only testing static parameters but also simulating the satellite's actual workload and operating conditions (such as high-speed data transmission, complex algorithm execution, and dynamic power switching) while under irradiation. This allows for a more realistic exposure of systemic faults or interaction problems caused by radiation under complex operating conditions.
[0233] Multi-factor coupling testing: Radiation testing is performed simultaneously or sequentially in conjunction with other environmental stresses (such as extreme temperature cycling, vacuum, vibration). This verifies the true performance of components in a comprehensive space environment and reveals the synergistic effects between stresses (e.g., high temperatures may exacerbate leakage current caused by radiation).
[0234] System-level / board-level radiation testing: This involves not only testing individual chips but also conducting overall radiation testing on key functional modules or subsystem circuit boards. This captures system-level effects (such as single-event transient propagation and synchronization failure) of inter-chip interfaces, power networks, and clock distribution under radiation.
[0235] Accelerated testing methods: Propose new acceleration factor models or test profile design methods that can more accurately predict long-term on-orbit radiation damage (TID effect) in a shorter test time, or more effectively cover the risk of high-energy particle-induced SEE (single event effect).
[0236] Advanced detectors and monitoring technologies: Utilizing detectors with ultra-high temporal / spatial resolution (such as transient current measurement and laser-induced single-event effect localization) or novel online monitoring technologies (such as built-in self-test and hardware Trojan detection technology for radiation effect capture) to achieve precise localization, rapid capture, and in-depth diagnosis of radiation-induced faults.
[0237] Artificial intelligence / big data analytics: Applying machine learning algorithms to analyze massive amounts of irradiation test data (parameter drift, fault logs, image data) to automatically identify fault modes, predict failure thresholds, assess margins, and even optimize subsequent testing plans. For example, AI can predict the sensitivity of a specific circuit structure to a certain particle.
[0238] Novel radiation hardening design verification technology: adopts an innovative simulation-test hybrid verification method, or applies advanced radiation effect simulation tools (combined with TCAD, Spice, system-level models) to predict hotspots before testing, guide test design, and reproduce and interpret phenomena with high fidelity after testing.
[0239] Innovative applications of microbeam / focused beam technology: Using high-precision microbeams or focused ion beams, specific sensitive areas of chips (such as memory cells and sensitive nodes) are precisely irradiated in a "surgical" manner to quantitatively study local radiation effects and verify the effectiveness of hardening designs.
[0240] Simulation of more realistic / extreme environments: More powerful simulation facilities or methods have been developed to more accurately reproduce the radiation spectrum of specific mission orbits (such as the high radiation belts of Jupiter) or simulate the extreme flux conditions of super solar particle events.
[0241] Hybrid field simulation: Simulate multiple radiation effects (such as TID and SEE) simultaneously in a single test, or simulate the coupled environment of radiation and electromagnetic interference.
[0242] Data analysis and application of results:
[0243] Innovation in probabilistic risk assessment models: Based on test data, a more refined probabilistic physical model is established to predict the on-orbit failure rate of satellite systems, especially for complex SEEs (such as single-event failures) or cascading effects.
[0244] The innovative design feedback loop not only uses test results for "pass / fail" judgments, but also establishes an efficient and automated feedback mechanism that directly and quickly feeds back failure mechanisms and quantitative data to the design team for iterative optimization of circuit design, layout, or hardening strategies.
[0245] The establishment of new reliability / safety indicators: Based on innovative testing and analysis, quantitative indicators of radiation reliability or safety are defined to provide more scientific guidance for design and mission planning.
[0246] While some embodiments of the present invention have been described in this application, those skilled in the art will understand that these embodiments are merely illustrative. Numerous variations, alternatives, and improvements will arise in those skilled in the art under the teachings of this invention without departing from its scope. The appended claims are intended to define the scope of the invention and thereby cover methods and structures within the scope of the claims themselves and their equivalents.
Claims
1. A method of satellite security assessment, characterized by, Comprise: determining the evaluation scope and target; collecting information of the satellite system; based on the evaluation scope and target, performing preliminary system safety evaluation according to the information of the satellite system, determining the satellite system safety requirement at the initial stage of satellite system design, and distributing the satellite system safety requirement to each subsystem of the satellite system, and preliminarily analyzing potential hazards and impact degree; and based on the evaluation scope and target, performing detailed system safety evaluation according to the information of the satellite system, further analyzing system failure mode and impact at the detailed design stage, determining risk level and failure probability, and quantitatively analyzing whether the safety requirement is met.
2. The satellite safety assessment method according to claim 1, characterized by, The determination of the evaluation scope and target comprises: the evaluation scope covers the whole life cycle of the satellite from launch to end of life, including the satellite body and ground support system; the evaluation target is to ensure that the safety risk of the satellite is acceptable under the specified mission scenario, and to meet the relevant standards and user requirements; and the collection of information of the satellite system comprises collecting design documents and operation data, wherein: the design documents comprise satellite overall design scheme, subsystem design drawings and technical specifications to clearly define the satellite system architecture, function and performance index; the operation data comprise satellite on-orbit operation history data for analyzing problems and potential risks in actual operation.
3. The satellite safety assessment method of claim 1, wherein, The preliminary system safety evaluation comprises: identifying potential hazards according to the satellite mission and operating environment; classifying the hazards according to severity; and distributing the overall safety requirement to each subsystem.
4. The satellite safety assessment method according to claim 3, characterized by, The identification of potential hazards according to the satellite mission and operating environment comprises: performing functional hazard analysis to analyze the hazards caused by functional failure one by one; identifying single point failure and hazards caused by single point failure; performing preliminary hazard analysis to identify top-level hazards based on mission scenarios; performing event tree analysis to deduce the consequences caused by initial events; performing specific space environment analysis, including space debris collision probability model analysis of collision probability, radiation effect analysis, and surface charging and discharge risk analysis caused by magnetic storm.
5. The method of claim 3, wherein, The classification of hazards according to severity comprises: classifying the hazards in order of decreasing severity, including catastrophic hazard, dangerous hazard, major hazard, minor hazard and negligible hazard, wherein: the catastrophic hazard is the hazard that causes death of personnel or complete loss of system; the dangerous hazard is the hazard that causes serious injury of personnel or serious damage of system; the major hazard is the hazard that causes significant degradation of mission capability; the minor hazard is the hazard that causes reduction of mission efficiency but can be recovered; and the negligible hazard is the hazard that has no safety impact; The distribution of the overall safety requirement to each subsystem comprises: decomposing the system-level safety target to each subsystem based on quantitative distribution and qualitative requirement, and determining the safety target of each subsystem, wherein the quantitative distribution is based on fault tree analysis to calculate the contribution degree of each subsystem, and multiplying the contribution degree by the system-level safety target to determine the safety target of each subsystem.
6. The method of claim 1, wherein, The detailed system safety evaluation comprises failure mode and effect analysis and fault tree analysis, wherein the failure mode and effect analysis comprises: performing failure mode and effect analysis on each subsystem, component and software of the satellite to determine the impact, possible cause and detection method of each failure mode; According to the severity of the failure impact and the probability of occurrence, the risk priority number RPN is calculated.
7. The method of claim 6, wherein, The calculation method of RPN is: RPN = SEV × OCC × DET, wherein: SEV represents severity, which is used to evaluate the degree of harm to system, task or personnel safety after failure, usually divided into 1-10 levels, 10 levels are the most serious; OCC occurrence degree, which is used to evaluate the probability of failure cause, usually divided into 1-10 levels, 10 levels are the most frequent; DET represents the detection degree, which is used to evaluate the ability of existing control measures to detect the cause or failure mode before the failure occurs, divided into 1-10 levels, 10 levels are almost impossible to detect.
8. The method of claim 6, wherein, Fault tree analysis includes building a fault tree and analyzing a fault tree, wherein: Building a fault tree includes: taking a major hazard event as the top event, combining the logical relationship of the satellite system, and building a fault tree; Analyzing the fault tree includes: qualitative analysis to find out the minimum cut set and determine the key failure mode, including: Using the down method or the up method to traverse the fault tree, finding out all the basic event sets that can cause the top event to occur to obtain the cut set, a cut set is a set of basic events, if all events in the set occur, the top event will inevitably occur; All cut sets obtained are minimized, wherein a cut set is a minimum cut set, and only when any basic event in the cut set is removed, the remaining set is no longer a cut set, that is, it can no longer cause the top event to occur alone; Sort the minimum cut sets, including: Ordering: first-order minimum cut set > second-order minimum cut set > third-order minimum cut set >... The lower the order, the higher the risk; Within the first-order minimum cut set, or between minimum cut sets of the same order, sort the basic events according to their failure rates, severity, detectability, and maintainability; According to the sorting results of the minimum cut sets, identify the key failure modes, improve the design, evaluate the effectiveness of redundancy, and develop maintenance and testing strategies.
9. The method of claim 8, wherein, When performing quantitative analysis, input the basic event failure rate data λ, task time t, common cause failure parameter β, and maintenance strategy; Quantitative analysis includes: Calculate the probability of basic events, including: usually assume constant failure rate, the probability of basic event failure within task time t P(Basic) ≈ λ * t; Calculate the top event probability using the minimum cut set MCS, including: The probability of the top event occurring is approximately equal to the union of the probabilities of all minimum cut sets, wherein the probability of the top event P(Top) ≈ ΣP(MCSi)-ΣΣP(MCSi∩MCSj)+ΣΣΣP(MCSi∩MCSj∩MCSk)-...; When calculating the probability of a minimum cut set containing redundant components, the probability of a common cause failure event needs to be added to the calculation of the corresponding minimum cut set, wherein the probability of a common cause failure event is calculated based on the common cause failure parameter β; External events are treated as independent basic events, and their occurrence probabilities are involved in the calculation.
10. The method of claim 1, wherein, Also includes: Verify whether the satellite system design and implementation meet the safety requirements, and confirm whether the safety of the satellite system meets the requirements in actual use.
Citation Information
Cited By
Space debris risk assessment method and system based on multi-dimensional index fusion
CN122508864A