A data security protection method and system for civil aviation data middleware
By distributing and storing highly confidential data and generating obfuscated variables, combined with correlation blocking methods and backup strategies, the risk of network attackers computing highly confidential data is mitigated, thereby improving data security and fully utilizing data value.
Patent Information
- Application Number
- CN202511493542.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-20
AI Technical Summary
How can we improve data security by reducing the risk that cyber attackers can calculate higher-secrecy data using correlation features between lower-secrecy and higher-secrecy data, while fully leveraging the data value of correlation features?
High-security data is divided into N data sub-blocks according to data type and stored in N storage modules. A random number generation algorithm is used to generate obfuscated variables, which are superimposed with the first associated data to obtain the second associated data. The obfuscated variables are then decomposed into N obfuscated sub-blocks and stored in a distributed manner. A restricted access policy is generated through an association blocking method, and the obfuscated sub-blocks are backed up after a formatting command to ensure data security.
It effectively prevents network attackers from deducing high-security data by cracking low-security data, while allowing legitimate users to aggregate obfuscated sub-blocks to obtain data value, thereby improving data security and calculating the estimated value of obfuscated variables under fault-tolerant conditions.
Smart Images

Figure CN120951402B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of civil aviation data security technology, specifically a data security protection method and system for a civil aviation data platform. Background Technology
[0002] The Civil Aviation Data Platform is a core platform for integrating, managing, and analyzing multi-source data in the civil aviation industry. Its data covers the entire lifecycle and all business scenarios of civil aviation operations, including core business operation data, customer and passenger data, flight and safety data, logistics and cargo data, external environmental data, and aircraft sensor data. The confidentiality levels of these data vary. For example, aircraft sensor data involves aircraft performance parameters and has a high confidentiality level. External environmental data, including airspace weather and environmental data, has a lower confidentiality level.
[0003] With the development of big data technology, technicians often need to rely on civil aviation data platforms to correlate and mine multi-source data to obtain variables that reflect the inherent correlation characteristics of the data, which are then recorded as correlation feature data. However, this process involves data security issues. For example, if technicians combine external environmental data and aircraft sensor data and mine the correlation between the two, then the aircraft sensor data is reflected in the correlation between the external environmental data and the aircraft sensor data. Even if security management personnel regularly destroy the aircraft sensor data, cyber attackers can still calculate the highly confidential aircraft sensor data by using the correlation between the external environmental data and the aircraft sensor data, and by analyzing the external environmental data with lower security levels.
[0004] Therefore, how to improve data security while fully leveraging the data value of correlation feature data, and reducing the risk that network attackers can calculate high-security data through low-security data and the correlation feature data between low-security and high-security data, is a technical problem that needs to be solved. Summary of the Invention
[0005] (1) Technical problems to be solved
[0006] The purpose of this invention is to provide a data security protection method and system for a civil aviation data platform, so as to reduce the risk of network attackers calculating high-secrecy-level data through low-secrecy-level data and the correlation characteristic data between low-secrecy-level data and high-secrecy-level data, while giving full play to the data value of correlation characteristic data, thereby improving data security.
[0007] (2) Technical solution
[0008] To achieve the above objectives, the present invention provides a data security protection method for a civil aviation data platform, the method comprising the following steps:
[0009] S1, Obtain first data; the first data includes civil aviation multi-source data and correlation feature data between civil aviation multi-source data; divide the civil aviation multi-source data into high-security level data and low-security level data; record the correlation feature data between high-security level data and low-security level data in the correlation feature data between civil aviation multi-source data as the first correlation data.
[0010] S2, the high-security data is divided according to data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module; a random number generation algorithm is used to generate obfuscated variables; the obfuscated variables are superimposed with the first associated data to obtain the second associated data; the obfuscated variables are decomposed into... N A confusing sub-block; the aforementioned N The obfuscated sub-blocks are stored separately in N In each storage module; the obfuscation variable is destroyed; N One storage module, the second associated data, and low-security-level data are open to user access.
[0011] S3, Based on the second associated data and the low-security-level data, a restricted access policy is generated using an association blocking method; the association blocking method triggers an abnormal alarm for requests that simultaneously access the low-security-level data and the second associated data.
[0012] S4, upon receiving a formatting command, first back up the obfuscated sub-blocks stored in the storage module that received the formatting command to the storage module that did not receive the formatting command before formatting.
[0013] Furthermore, the correlation characteristics between the multi-source civil aviation data are obtained in advance through data mining of the multi-source civil aviation data, representing the correlation between the multi-source civil aviation data.
[0014] Furthermore, the method for dividing the multi-source civil aviation data into high-security-level data and low-security-level data includes:
[0015] Based on the pre-obtained data security levels of multi-source civil aviation data, the multi-source civil aviation data is divided into data of security levels from Level 1 to Level 2. Q Confidentiality level data; Q This indicates the number of data confidentiality levels.
[0016] The first Security level data up to the QClassified data is designated as high-class data; data classified as first-class data is then classified as high-class data. Data with a security classification level is recorded as low-security data; among which This indicates a pre-set confidentiality threshold.
[0017] Furthermore, the aforementioned N Each storage module is connected via a distributed network; It is a pre-set positive integer greater than 2.
[0018] Furthermore, the method for superimposing the confusion variable with the first association data to obtain the second association data includes:
[0019] The second association data is calculated using a superposition formula based on the confounding variable and the first association data; the superposition formula is as follows:
[0020] ;
[0021] in, Indicates the first related data. This indicates the second related data. This indicates a confusing variable.
[0022] Furthermore, the decomposition of the confusion variable into N The methods for obfuscating sub-blocks include:
[0023] Generate using hash function method A set of independent random numbers, denoted as the first independent random number. To the Independent random numbers The The range of values for each independent random number is within arrive Between; among This represents a pre-set error control constant with a value greater than 0.
[0024] According to the above The relevant perturbation is obtained by calculating a number of independent random numbers. ; The calculation formula is:
[0025] ;
[0026] in, Indicates the first Independent random numbers; For values from 1 to Positive integers between [a certain range].
[0027] In sequence to , cloned as to ;Will to These are respectively denoted as the first offset to the second offset. Offset.
[0028] Based on the aforementioned confusion variable, the first offset to the... Offset calculation obtained N There are 1 to 10 obfuscation sub-blocks, respectively denoted as the first obfuscation sub-block to the second obfuscation sub-block. N Obfuscated sub-blocks; where the first The formula for calculating the obfuscated sub-block is:
[0029] ;
[0030] in, Indicates the first Obfuscate sub-blocks; For values from 1 to Positive integers between [a certain range].
[0031] Furthermore, the method of backing up the obfuscated sub-block stored in the storage module that received the formatting instruction to the storage module that did not receive the formatting instruction before formatting after receiving the formatting instruction includes:
[0032] Regarding the N Each storage module is numbered. N The storage modules are respectively referred to as the first storage module to the sixth storage module. N Storage module; search the first storage module to the second. N The connection topology of the storage modules in the distributed network; the first storage module to the second... N The storage modules are connected via a network link; respectively, the first storage module to the second storage module are connected via a network link. N Storage modules directly connected via network links are designated as the first backup storage module set. N A collection of backup storage modules.
[0033] Searching from the first storage module to the second storage module at a preset first interval time. N The storage module sends a formatting command; the storage module that receives the formatting command is marked as a storage module to be destroyed; the globally optimal backup storage module for the storage module to be destroyed is selected from the set of backup storage modules corresponding to the storage module to be destroyed; the obfuscated sub-blocks stored in the storage module to be destroyed are backed up to the corresponding globally optimal backup storage module; after a second interval, the storage module to be destroyed is formatted; if the globally optimal backup storage module does not exist, the storage module to be destroyed is formatted after a second interval.
[0034] Furthermore, the method for selecting the globally optimal backup storage module from the set of backup storage modules corresponding to the storage module to be destroyed includes:
[0035] The storage modules to be destroyed are renumbered from the first storage module to the second. H Storage modules to be destroyed; among them H The number of storage modules to be destroyed; filter from the first storage module to the second. H The storage modules in the backup storage module set corresponding to the storage module to be destroyed that have not received a formatting command are respectively denoted as the first set to the second set. H Set; evaluate the first storage module to be destroyed up to the second respectively. H Storage modules to be destroyed and the first set to the second set H The network transmission performance metrics of the network links between storage modules in the set; the network transmission performance metrics are packet loss rate; the search yields the first set to the... H The data security level corresponding to the data sub-blocks stored in the storage modules of the collection.
[0036] According to the first storage module to be destroyed to the second H Storage modules to be destroyed and the first set to the second set H Network transmission performance metrics of network links between storage modules in the set and the first set to the second set H The data security level corresponding to the data sub-blocks stored in the storage modules of the set is used to construct an optimization objective function; the independent variable of the optimization objective function is the number of the first candidate backup storage module up to the [number missing]. H The numbers of the candidate backup storage modules; wherein, the first candidate backup storage module to the second... H The value ranges for the candidate backup storage modules are from the first set to the second set. H gather.
[0037] With the objective function being minimized, the particle swarm optimization algorithm was used to calculate the first candidate backup storage module numbering up to the [number missing]. H The optimal values for the candidate backup storage module numbers are denoted as the first optimal number to the second optimal number. H Optimal numbering; assign the first optimal number to the next optimal number. H The storage module corresponding to the optimal number is denoted as the first globally optimal backup storage module up to the next. H Globally optimal backup storage module.
[0038] Furthermore, the optimization objective function is:
[0039] ;
[0040] in, This represents the objective function to be optimized. Indicates the first The number of the backup storage module to be selected. The value is 1 to Integers between [a certain number] For the first Storage modules to be destroyed and the first Packet loss rate of network links between storage modules To determine the data security level and data storage period based on the pre-set mapping relationship, the first... After the data security level corresponding to the data sub-blocks stored in the storage module is converted into the data storage cycle, it is then processed through a pre-set reverse mapping relationship and normalized to a value within the range of 0 to 1. For the pre-set first weight, The second weight is set in advance; and The sum of is 1.
[0041] Based on the same inventive concept, this invention also provides a data security protection system for a civil aviation data platform, the system comprising, in sequence: a first data acquisition module, a data obfuscation module, an association blocking module, and a formatting module.
[0042] The first data acquisition module is used to acquire first data; the first data includes civil aviation multi-source data and correlation feature data between civil aviation multi-source data; the civil aviation multi-source data is divided into high-security level data and low-security level data; the correlation feature data between high-security level data and low-security level data in the correlation feature data between civil aviation multi-source data is recorded as the first correlation data.
[0043] The data obfuscation module is used to divide the high-security data according to its data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module; a random number generation algorithm is used to generate obfuscated variables; the obfuscated variables are superimposed with the first associated data to obtain the second associated data; the obfuscated variables are decomposed into... N A confusing sub-block; the aforementioned N The obfuscated sub-blocks are stored separately in N In each storage module; the obfuscation variable is destroyed; N One storage module, the second associated data, and low-security-level data are open to user access.
[0044] The correlation blocking module is used to generate a restricted access policy based on the second correlation data and the low-security-level data using a correlation blocking method; the correlation blocking method is to trigger an abnormal alarm for requests that simultaneously access the low-security-level data and the second correlation data.
[0045] The formatting module is used to, upon receiving a formatting instruction, first back up the obfuscated sub-block stored in the storage module that received the formatting instruction to the storage module that did not receive the formatting instruction before performing the formatting.
[0046] (3) Beneficial effects
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] 1. Classify high-security data according to data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module. Then, a random number generation algorithm is used to generate obfuscated variables, which are then superimposed with the first associated data to obtain the second associated data. The obfuscated variables are decomposed into N A confusing sub-block, which will... N The obfuscated sub-blocks are stored separately in N The obfuscated variables are then destroyed within the storage module. N The first storage module, the second associated data, and the low-security-level data are open to user access. This approach serves two purposes: firstly, it prevents network attackers from deducing the high-security-level data by cracking the low-security-level data and the first associated data; secondly, it allows researchers who genuinely need to use the first associated data, high-security-level data, and low-security-level data for data analysis to access these resources. N The obfuscated sub-blocks in each storage module are aggregated to obtain the first related data and the high-security-level data. This method enhances data security while fully leveraging the value of the related data.
[0049] 2. The method of decomposing obfuscated variables takes into account a certain degree of fault tolerance. That is, even if the network connection of one or more storage modules fails, or if there is no network link between a storage module and other unformatted storage modules, the estimated value of the obfuscated variable can still be calculated based on the available obfuscated sub-blocks, and thus the estimated value of the first associated data can be obtained. In this case, the estimated value of the obfuscated variable has an error, but the range of the error can be calculated based on the number of lost obfuscated sub-blocks.
[0050] 3. Construct an optimization objective function and use the particle swarm optimization algorithm to solve it, thereby obtaining the optimal obfuscated sub-block backup strategy. This backup strategy comprehensively considers network transmission performance indicators and the expected number of backups. Attached Figure Description
[0051] Figure 1 This is a flowchart of a data security protection method for a civil aviation data platform according to Embodiment 1 of the present invention;
[0052] Figure 2 This is a schematic diagram of the module composition of a data security protection system for a civil aviation data platform according to Embodiment 2 of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Before providing examples, it's necessary to explain the application scenario of this invention, which is applied to the security protection of a civil aviation data platform. The civil aviation data platform is a core platform for integrating, managing, and analyzing multi-source data in the civil aviation industry. Its data covers the entire lifecycle and all business scenarios of civil aviation operations, including core business operation data, customer and passenger data, flight and safety data, logistics and cargo data, external environmental data, and aircraft sensor data. These data have varying levels of confidentiality. For example, aircraft sensor data involves aircraft performance parameters and has a high level of confidentiality. External environmental data, including airspace weather and environmental data, has a lower level of confidentiality. In practice, technicians often need to rely on the civil aviation data platform to correlate and mine multi-source data to obtain variables that reflect the inherent correlation characteristics of the data. For example, aircraft sensor data includes aircraft engine vibration amplitude data; external environmental data includes average wind speed data, turbulence intensity data, etc. Turbulence intensity data represents the ratio of the standard deviation of wind speed fluctuations to the average wind speed data, reflecting the degree of airflow turbulence. Data mining revealed a linear positive correlation between aircraft engine vibration amplitude data and the product of turbulence intensity data and the square of average wind speed data. In other words, the ratio of the change in aircraft engine vibration amplitude data to the change in turbulence intensity data multiplied by the square of average wind speed data is approximately equal to a constant. This constant is denoted as the vibration amplitude proportionality coefficient. This coefficient is related to factors such as damping ratio and the load-bearing area of the aircraft blades. Assuming researchers have already obtained the vibration amplitude proportionality coefficient for a specific aircraft through data mining, further in-depth research can be conducted in conjunction with this coefficient, aircraft engine vibration amplitude data, average wind speed data, and turbulence intensity data during internal work on civil aviation data. However, these data have different levels of confidentiality. Aircraft engine vibration amplitude data involves aircraft performance parameters and is classified as high-level confidentiality data. Average wind speed data and turbulence intensity data are external environmental data and have lower confidentiality levels. The vibration amplitude proportionality coefficient is a variable derived from existing data that reflects the inherent correlation characteristics of the data. Because new variables reflecting the inherent correlations within the data are constantly discovered during the research process, it is impossible to classify their security levels using existing, fixed security classification systems. However, this raises a new problem: aircraft engine vibration amplitude data, due to its high security level, requires periodic destruction. Average wind speed data and turbulence intensity data, with lower security levels, do not require periodic destruction. However, the security management of vibration amplitude ratio coefficients remains a gap. If vibration amplitude ratio coefficients are not destroyed, cyber attackers can use them and low-security wind speed and turbulence intensity data to deduce aircraft engine vibration amplitude data, leading to the leakage of high-security data.However, if the vibration amplitude ratio coefficient is destroyed along with the aircraft engine vibration amplitude data, the lifespan of the data's value is greatly reduced, and the data value of the vibration amplitude ratio coefficient, which was calculated with considerable effort, is wasted. The purpose of this embodiment is to reduce the risk of network attackers calculating high-security data from low-security data and the correlation characteristics between low-security and high-security data, while fully leveraging the data value of correlation feature data, thereby improving data security.
[0055] Example 1: As Figure 1 As shown in the figure, this embodiment provides a data security protection method for a civil aviation data platform, the method including the following steps:
[0056] S1, Obtain first data; the first data includes civil aviation multi-source data and correlation feature data between civil aviation multi-source data; divide the civil aviation multi-source data into high-security level data and low-security level data; record the correlation feature data between high-security level data and low-security level data in the correlation feature data between civil aviation multi-source data as the first correlation data.
[0057] For example, first data is acquired. This first data includes multi-source civil aviation data and correlation characteristic data between these multi-source data. The multi-source civil aviation data includes core business operation data, customer and passenger data, flight and safety data, logistics and freight data, external environment data, and aircraft sensor data; wherein, logistics and freight data and external environment data are low-level confidentiality data, while core business operation data, customer and passenger data, flight and safety data, and aircraft sensor data are high-level confidentiality data. The correlation characteristic data between the multi-source civil aviation data represents the relationships between these data, such as the correlation coefficient (denoted as the transportation time ratio coefficient) and vibration amplitude ratio coefficient between the average cargo transportation time data in logistics and freight data and the average wind speed data in external environment data. The transportation time ratio coefficient reflects the correlation characteristic data between low-level confidentiality data, therefore, the application of the transportation time ratio coefficient does not present a contradiction between data lifecycle and data security. The vibration amplitude ratio coefficient reflects the relationship between aircraft engine vibration amplitude data and average wind speed data and turbulence intensity data. Aircraft engine vibration amplitude data belongs to aircraft sensor data and is classified as high-security data. Average wind speed data and turbulence intensity data belong to external environmental data and are classified as low-security data. Therefore, the vibration amplitude proportionality coefficient, representing the correlation characteristic between high-security and low-security data, is denoted as the first correlation data and has a value of 1.28. .
[0058] S2, the high-security data is divided according to data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module; a random number generation algorithm is used to generate obfuscated variables; the obfuscated variables are superimposed with the first associated data to obtain the second associated data; the obfuscated variables are decomposed into... N A confusing sub-block; the aforementioned N The obfuscated sub-blocks are stored separately in N In each storage module; the obfuscation variable is destroyed; N One storage module, the second associated data, and low-security-level data are open to user access.
[0059] For example, the highly confidential data is divided into 96 data sub-blocks according to data type, denoted as storage modules 1 to 96. These 96 data sub-blocks are distributed across the 96 storage modules. These 96 storage modules are located in eight data centers at different geographical locations, and each data center stores data containing all data types. For instance, storage modules 1 to 12 are located in data center 1, where the data sub-blocks stored in storage modules 1 to 3 contain core business operation data, storage modules 4 to 6 contain customer and passenger data, storage modules 7 to 9 contain flight and safety data, and storage modules 10 to 12 contain aircraft sensor data. And so on. A random number generation algorithm is used to generate obfuscated variables, the generation period of which can be manually set, and the values of these obfuscated variables range from -200 to 200. For example, in this embodiment, since the longest data storage period for all high-security data is 30 days, the generation period for the obfuscation variable is also set to 30 days, meaning the obfuscation variable is generated every 30 days. The resulting obfuscation variable is 136. Based on the obfuscation variable and the first associated data 1.28, the second associated data is calculated to be 3.0208. The obfuscation variable 136 is decomposed into 96 obfuscation sub-blocks, and these 96 obfuscation sub-blocks are distributed and stored in 96 storage modules. The 96 obfuscation sub-blocks are concatenated with the 96 data sub-blocks and then encrypted. The encryption algorithm is denoted as the first encryption algorithm. The obfuscation variable is erased in the server storing the obfuscation variable, thereby destroying the obfuscation variable and preventing users from directly obtaining it in any way. The 96 storage modules, the second associated data (i.e., 3.0208), and the low-security data are made available to users. For users who need to use the first associated data, access permissions to the 96 storage modules, the second associated data (i.e., 3.0208), and the low-security data are granted, and the decryption algorithm corresponding to the first encryption algorithm is transmitted to the user. For users with lower privileges, who are only allowed access to low-security-level data, access is only granted to that data. The 96 storage modules store not only high-security-level data but also 96 obfuscated sub-blocks. Although the obfuscated variables have been destroyed, once the 96 storage modules are open to users, they can decrypt the encrypted data within these modules, extract and aggregate the 96 obfuscated sub-blocks to calculate an estimated value for the obfuscated variables, denoted as the obfuscation estimate. This estimated value is then subtracted from the obfuscated estimate to obtain the estimated value for the first associated data. Using this method, users who need the first associated data can obtain the data value of high-security-level data, low-security-level data, and the first associated data, while simultaneously improving data security.
[0060] S3, Based on the second associated data and the low-security-level data, a restricted access policy is generated using an association blocking method; the association blocking method triggers an abnormal alarm for requests that simultaneously access the low-security-level data and the second associated data.
[0061] For example, when a user or cyber attacker simultaneously requests access to low-security-level data and second-related data within a pre-defined time interval, it may create a risk of inferring high-security-level data from the low-security-level data. This would trigger an anomaly alert, which would be sent to the administrator. The administrator would then assess the access behavior and decide whether to allow data access.
[0062] S4, upon receiving a formatting command, first back up the obfuscated sub-blocks stored in the storage module that received the formatting command to the storage module that did not receive the formatting command before formatting.
[0063] Because high-security data needs to be destroyed through periodic formatting, 96 storage modules will be formatted one by one over 30 days. However, these 96 modules store not only high-security data but also 96 obfuscated sub-blocks. If left unattended, these obfuscated sub-blocks will be destroyed along with the high-security data. However, if only the high-security data is formatted during the module formatting process, while the obfuscated sub-blocks are retained, it will waste storage space and maintenance costs. This is because although obfuscated sub-blocks occupy very little storage space, retaining them after formatting the high-security data requires power supply and maintenance to the storage modules, consuming storage space for data indexing and connections, even if the amount of data stored in the modules is small, resulting in unnecessary waste. Therefore, upon receiving a formatting command, the obfuscated sub-blocks stored in the modules receiving the command are first extracted using a decryption algorithm. These obfuscated sub-blocks are then backed up to the modules that have not received the command before formatting. Using this method, as long as all 96 storage modules are not yet formatted, all information in the 96 obfuscated sub-blocks is preserved, and the storage capacity of the storage modules is fully utilized. However, if all 96 storage modules have been formatted, it indicates that the longest data storage period for high-security data has been reached, and the obfuscated sub-blocks will be completely destroyed. The data value of the first associated data has already been fully utilized within this longest data storage period. From then on, cyber attackers will be unable to deduce the high-security data from the second associated data or the low-security data in any way, thus ensuring data security.
[0064] Furthermore, the correlation characteristics between the multi-source civil aviation data are obtained in advance through data mining of the multi-source civil aviation data, representing the correlation between the multi-source civil aviation data.
[0065] Furthermore, the method for dividing the multi-source civil aviation data into high-security-level data and low-security-level data includes:
[0066] Based on the pre-obtained data security levels of multi-source civil aviation data, the multi-source civil aviation data is divided into data of security levels from Level 1 to Level 2. Q Confidentiality level data; Q This indicates the number of data confidentiality levels.
[0067] The first Security level data up to the Q Classified data is designated as high-class data; data classified as first-class data is then classified as high-class data. Data with a security classification level is recorded as low-security data; among which This indicates a pre-set confidentiality threshold.
[0068] For example, the civil aviation multi-source data is divided into five levels of confidentiality: Level 1 to Level 5. Specifically, core business operation data is Level 3 confidentiality data, customer and passenger data is Level 3 confidentiality registration data, flight and safety data is Level 4 confidentiality data, logistics and cargo data is Level 2 confidentiality data, external environment data is Level 1 confidentiality data, and aircraft sensor data is Level 5 confidentiality data. A pre-set confidentiality threshold of three is used, meaning that Level 1 to Level 2 confidentiality data are considered low-confidentiality data, and Level 3 to Level 5 confidentiality data are considered high-confidentiality data. Therefore, logistics and cargo data and external environment data are low-confidentiality data, while core business operation data, customer and passenger data, flight and safety data, and aircraft sensor data are high-confidentiality data.
[0069] Furthermore, the aforementioned N Each storage module is connected via a distributed network; It is a pre-set positive integer greater than 2.
[0070] For example, 96 storage modules are connected via a distributed wireless network. Distributing data storage can reduce the risk of data breaches caused by single-point attacks from cyber attackers.
[0071] Furthermore, the method for superimposing the confusion variable with the first association data to obtain the second association data includes:
[0072] The second association data is calculated using a superposition formula based on the confounding variable and the first association data; the superposition formula is as follows:
[0073] ;
[0074] in, Indicates the first related data. This indicates the second related data. This indicates a confusing variable.
[0075] For example, the confusion variable is 136, the first associated data is 1.28, and therefore the second associated data is calculated to be 3.0208.
[0076] Furthermore, the decomposition of the confusion variable into N The methods for obfuscating sub-blocks include:
[0077] Generate using hash function method A set of independent random numbers, denoted as the first independent random number. To the Independent random numbers The The range of values for each independent random number is within arrive Between; among This represents a pre-set error control constant with a value greater than 0.
[0078] According to the above The relevant perturbation is obtained by calculating a number of independent random numbers. ; The calculation formula is:
[0079] ;
[0080] in, Indicates the first Independent random numbers; For values from 1 to Positive integers between [a certain range].
[0081] In sequence to , cloned as to ;Will to These are respectively denoted as the first offset to the second offset. Offset.
[0082] Based on the aforementioned confusion variable, the first offset to the... Offset calculation obtained N There are 1 to 10 obfuscation sub-blocks, respectively denoted as the first obfuscation sub-block to the second obfuscation sub-block. N The j-th obfuscated sub-block is calculated using the following formula:
[0083] ;
[0084] in, This represents the j-th obfuscated sub-block; j is a value ranging from 1 to... Positive integers between [a certain range].
[0085] For example, a hash function is used to generate 95 independent random numbers with values between -2 and 2, which are denoted as the first independent random number. Up to the 95th independent random number The error control constant is set to 2. The relevant disturbances are calculated based on the first to the 95th independent random numbers. .Will to Cloning yields offsets from the first to the 96th, resulting in the first to the 96th obfuscated sub-blocks. This technique offers three main advantages: First, the randomly generated obfuscated sub-blocks enhance data security. Second, ideally, the sum of the obfuscated sub-blocks is always equal to the obfuscation variable. This means that if a user can access all obfuscated sub-blocks, they can calculate the obfuscation variable by summing them. Therefore, even after the obfuscation variable is destroyed, the user can still obtain an estimated value of the obfuscation variable if they have access to all obfuscated sub-blocks. Third, considering special cases, such as network connectivity failures in one or more storage modules, or the lack of network link between a storage module and other unformatted storage modules, the user cannot obtain all obfuscated sub-blocks. The method proposed in this embodiment allows the user to calculate an estimated value of the obfuscation variable based on the available obfuscated sub-blocks. In such cases, the estimated value of the obfuscation variable may have errors, but the range of these errors can be calculated based on the number of lost obfuscated sub-blocks.
[0086] Furthermore, the method of backing up the obfuscated sub-block stored in the storage module that received the formatting instruction to the storage module that did not receive the formatting instruction before formatting after receiving the formatting instruction includes:
[0087] Regarding the N Each storage module is numbered. N The storage modules are respectively referred to as the first storage module to the sixth storage module. N Storage module; search the first storage module to the second. N The connection topology of the storage modules in the distributed network; the first storage module to the second... N The storage modules are connected via a network link; respectively, the first storage module to the second storage module are connected via a network link. N Storage modules directly connected via network links are designated as the first backup storage module set. N A collection of backup storage modules.
[0088] Searching from the first storage module to the second storage module at a preset first interval time. N The storage module sends a formatting command; the storage module that receives the formatting command is marked as a storage module to be destroyed; the globally optimal backup storage module for the storage module to be destroyed is selected from the set of backup storage modules corresponding to the storage module to be destroyed; the obfuscated sub-blocks stored in the storage module to be destroyed are backed up to the corresponding globally optimal backup storage module; after a second interval, the storage module to be destroyed is formatted; if the globally optimal backup storage module does not exist, the storage module to be destroyed is formatted after a second interval.
[0089] For example, the connection topology of storage modules 1 through 96 in a distributed network is searched. Storage modules 1 through 96 are connected via network links. Storage modules directly connected to storage modules 1 through 96 via network links are denoted as the first backup storage module set to the 96th backup storage module set. Taking the first backup storage module set as an example, the search reveals that the first storage module is directly connected to the second, third, fourth, eighth, 12th, 13th, 25th, and 68th storage modules via network links. Therefore, the first backup storage module set is: {second storage module, third storage module, fourth storage module, eighth storage module, 12th storage module, 13th storage module, 25th storage module, 68th storage module}. The second storage module is directly connected to the first and third storage modules via network links. Therefore, the second backup storage module set is: {first storage module, third storage module}. The third storage module is directly connected to the first, second, 42nd, and 68th storage modules via a network link. Therefore, the third backup storage module set is: {first storage module, second storage module, 42nd storage module, 68th storage module}. Formatting commands sent to the first through 96th storage modules are searched at preset intervals (e.g., one hour). It is found that the first, second, and third storage modules received formatting commands, and therefore, they are marked as storage modules to be destroyed. The globally optimal backup storage module for each storage module to be destroyed is selected from the set of backup storage modules corresponding to these modules. The globally optimal backup storage module for the first storage module is found to be the 25th storage module. The globally optimal backup storage module for the second storage module does not exist, and the globally optimal backup storage module for the third storage module is the 68th storage module. Therefore, the first obfuscated sub-block stored in the first storage module is backed up to the 25th storage module, and the third obfuscated sub-block stored in the third storage module is backed up to the 68th storage module. After a second delay (1 minute), the first, second, and third storage modules are formatted. At this point, the obfuscated sub-blocks in the first and third storage modules are retained, while the obfuscated sub-blocks in the second storage module are lost. Users can calculate estimated values of the obfuscated variables based on the obfuscated sub-blocks stored in storage modules four through ninety-six. However, because the second obfuscated sub-block is lost, the estimated values of the obfuscated variables will have errors, the range of which can be calculated.
[0090] Furthermore, the method for selecting the globally optimal backup storage module from the set of backup storage modules corresponding to the storage module to be destroyed includes:
[0091] The storage modules to be destroyed are renumbered from the first storage module to the second. H Storage modules to be destroyed; among them H The number of storage modules to be destroyed; filter from the first storage module to the second. H The storage modules in the backup storage module set corresponding to the storage module to be destroyed that have not received a formatting command are respectively denoted as the first set to the second set. H Set; evaluate the first storage module to be destroyed up to the second respectively. H Storage modules to be destroyed and the first set to the second set H The network transmission performance metrics of the network links between storage modules in the set; the network transmission performance metrics are packet loss rate; the search yields the first set to the... H The data security level corresponding to the data sub-blocks stored in the storage modules of the collection.
[0092] According to the first storage module to be destroyed to the second H Storage modules to be destroyed and the first set to the second set H Network transmission performance metrics of network links between storage modules in the set and the first set to the second set H The data security level corresponding to the data sub-blocks stored in the storage modules of the set is used to construct an optimization objective function; the independent variable of the optimization objective function is the number of the first candidate backup storage module up to the [number missing]. H The numbers of the candidate backup storage modules; wherein, the first candidate backup storage module to the second... H The value ranges for the candidate backup storage modules are from the first set to the second set. H gather.
[0093] With the objective function being minimized, the particle swarm optimization algorithm was used to calculate the first candidate backup storage module numbering up to the [number missing]. H The optimal values for the candidate backup storage module numbers are denoted as the first optimal number to the second optimal number. H Optimal numbering; assign the first optimal number to the next optimal number. H The storage module corresponding to the optimal number is denoted as the first globally optimal backup storage module up to the next. H Globally optimal backup storage module.
[0094] For example, the storage modules to be destroyed are renumbered as first to third storage modules. Storage modules that have not received a formatting command from the backup storage module sets corresponding to the first to third storage modules are categorized into sets one to three. The first set is: {fourth storage module, eighth storage module, 12th storage module, 13th storage module, 25th storage module, 68th storage module}, the second set is empty, and the third set is: {42nd storage module, 68th storage module}. An optimization objective function is constructed, and the particle swarm optimization algorithm is used to calculate that the first optimal number is 25, the second optimal number does not exist, and the third optimal number is 68. Therefore, the first globally optimal backup storage module is the 25th storage module, the second globally optimal backup storage module does not exist, and the third globally optimal backup storage module is the 68th storage module.
[0095] Furthermore, the optimization objective function is:
[0096] ;
[0097] in, This represents the objective function to be optimized. Indicates the first The number of the backup storage module to be selected. The value is 1 to Integers between [a certain number] For the first Storage modules to be destroyed and the first Packet loss rate of network links between storage modules To determine the data security level and data storage period based on the pre-set mapping relationship, the first... After the data security level corresponding to the data sub-blocks stored in the storage module is converted into the data storage cycle, it is then processed through a pre-set reverse mapping relationship and normalized to a value within the range of 0 to 1. For the pre-set first weight, The second weight is set in advance; and The sum of is 1.
[0098] For example, the optimization objective function reflects two aspects of information. First, it considers the network transmission performance index of the network link between the storage module to be destroyed and the storage modules in the backup storage module set that have not received formatting instructions, characterized by the packet loss rate; a lower packet loss rate indicates better network transmission performance. Second, it considers the duration for which the obfuscated sub-blocks can be retained without any further action after the storage module to be destroyed backs up the obfuscated sub-blocks to the storage modules in the backup storage module set that have not received formatting instructions. This duration is related to the data confidentiality level. The higher the data confidentiality level, the shorter the data storage cycle. Therefore, if the obfuscated sub-blocks in the storage module to be destroyed are backed up to storage modules with higher data confidentiality levels, then the obfuscated sub-blocks are more likely to be backed up or cleared in the next first interval. Therefore, the optimization objective function is a comprehensive objective function that considers both the network transmission performance index and the expected number of backups. , Based on historical experience and needs.
[0099] Example 2: Based on the same inventive concept, such as Figure 2 As shown in the figure, this embodiment also provides a data security protection system for a civil aviation data platform. The system includes, in sequence, a first data acquisition module, a data obfuscation module, an association blocking module, and a formatting module.
[0100] The first data acquisition module is used to acquire first data; the first data includes civil aviation multi-source data and correlation feature data between civil aviation multi-source data; the civil aviation multi-source data is divided into high-security level data and low-security level data; the correlation feature data between high-security level data and low-security level data in the correlation feature data between civil aviation multi-source data is recorded as the first correlation data.
[0101] The data obfuscation module is used to divide the high-security data according to its data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module; a random number generation algorithm is used to generate obfuscated variables; the obfuscated variables are superimposed with the first associated data to obtain the second associated data; the obfuscated variables are decomposed into... N A confusing sub-block; the aforementioned N The obfuscated sub-blocks are stored separately in N In each storage module; the obfuscation variable is destroyed; N One storage module, the second associated data, and low-security-level data are open to user access.
[0102] The correlation blocking module is used to generate a restricted access policy based on the second correlation data and the low-security-level data using a correlation blocking method; the correlation blocking method is to trigger an abnormal alarm for requests that simultaneously access the low-security-level data and the second correlation data.
[0103] The formatting module is used to, upon receiving a formatting instruction, first back up the obfuscated sub-block stored in the storage module that received the formatting instruction to the storage module that did not receive the formatting instruction before performing the formatting.
[0104] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0105] Finally, it should be noted that although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for data security protection of a civil aviation data center, characterized in that, The method comprises the following steps: S1, obtaining first data; the first data comprises civil aviation multi-source data, and correlation characteristic data between the civil aviation multi-source data; the civil aviation multi-source data is divided into high-security-level data and low-security-level data; the correlation characteristic data between the high-security-level data and the low-security-level data in the correlation characteristic data between the civil aviation multi-source data is recorded as first correlation data; S2, the high-security data is divided according to data type. N Each data sub-block will contain the... N The data sub-blocks are stored separately. N In each storage module; a random number generation algorithm is used to generate obfuscated variables; the obfuscated variables are superimposed with the first associated data to obtain the second associated data; the obfuscated variables are decomposed into... N A confusing sub-block; the aforementioned N The obfuscated sub-blocks are stored separately in N In each storage module; the obfuscation variable is destroyed; N One storage module, the second associated data, and low-security-level data are open to user access; S3, generating a restricted access strategy by a correlation blocking method according to the second correlation data and the low-security-level data; the correlation blocking method triggers an abnormal alarm for a request of simultaneously accessing the low-security-level data and the second correlation data; S4, after receiving a formatting instruction, backing up the confused sub-block stored in the storage module receiving the formatting instruction to the storage module not receiving the formatting instruction, and then formatting; The correlation characteristic data between the civil aviation multi-source data is obtained by data mining on the civil aviation multi-source data in advance, and represents the correlation between the civil aviation multi-source data; The method for dividing the civil aviation multi-source data into high-security-level data and low-security-level data comprises: According to a data security level of the civil aviation multi-source data obtained in advance, the civil aviation multi-source data is divided into first security level data to nth security level data Q security level data; Q represents the number of data security levels The first The first Q The first The first represents a pre-set security threshold The N storage modules are connected through a distributed network; is a positive integer greater than 2 set in advance; The method for superimposing the confused variable and the first correlation data to obtain the second correlation data comprises: The second correlation data is calculated according to the confused variable and the first correlation data by using a superposition formula; the superposition formula is: ; wherein, represents the first association data, represents the second association data, represents a confounding variable; the method of decomposing the obfuscated variable into N a plurality of obfuscated sub-blocks includes: Adopting hash function method to generate An independent random number, recorded as a first independent random number To the Independent random number ; the value range of the Independent random number is between To ; wherein Error control constant greater than 0 is set in advance. According to the above The related disturbance is calculated according to the above ; The calculation formula is: ; wherein, represents the independent random number; is a positive integer having a value of 1 to between 1 and 100. in succession to , cloned as to ; the to are denoted as first offset to nth offset, respectively; According to the confusion variable, the first offset to the offset calculation N confusion sub-blocks, respectively, the first confusion sub-block to the N confusion sub-block; wherein the calculation formula of the first confusion sub-block is: ; wherein, represents the confusion sub-block; is a positive integer having a value of 1 to .
2. The data security protection method for a civil aviation data hub according to claim 1, wherein, The method for, after receiving a formatting instruction, backing up the confused sub-block stored in the storage module receiving the formatting instruction to the storage module not receiving the formatting instruction, and then formatting comprises: Regarding the N Each storage module is numbered. N The storage modules are respectively referred to as the first storage module to the sixth storage module. N Storage module; search the first storage module to the second. N The connection topology of the storage modules in the distributed network; the first storage module to the second... N The storage modules are connected via a network link; respectively, the first storage module to the second storage module are connected via a network link. N Storage modules directly connected via network links are designated as the first backup storage module set. N Backup storage module collection; search for a storage module to which the first interval time is set as a time interval from the first storage module to the last storage module N The method further comprises: sending a formatting instruction to the storage modules; marking the storage module that receives the formatting instruction as a storage module to be destroyed; selecting a globally optimal backup storage module of the storage module to be destroyed from a backup storage module set corresponding to the storage module to be destroyed; and after backing up the confused sub-blocks stored in the storage module to be destroyed to the corresponding globally optimal backup storage module, formatting the storage module to be destroyed after a second interval time.
3. The data security protection method for a civil aviation data hub according to claim 2, wherein, The method for screening the globally optimal backup storage module of the to-be-destroyed storage module from the backup storage module set corresponding to the to-be-destroyed storage module comprises: The storage modules to be destroyed are renumbered from the first storage module to the second. H Storage modules to be destroyed; among them H The number of storage modules to be destroyed; filter from the first storage module to the second. H The storage modules in the backup storage module set corresponding to the storage module to be destroyed that have not received a formatting command are respectively denoted as the first set to the second set. H Set; evaluate the first storage module to be destroyed up to the second respectively. H Storage modules to be destroyed and the first set to the second set H The network transmission performance metrics of the network links between storage modules in the set; the network transmission performance metrics are packet loss rate; the search yields the first set to the... H The data security level corresponding to the data sub-blocks stored in the storage modules of the collection; According to the first storage module to be destroyed to the second H Storage modules to be destroyed and the first set to the second set H Network transmission performance metrics of network links between storage modules in the set and the first set to the second set H The data security level corresponding to the data sub-blocks stored in the storage modules of the set is used to construct an optimization objective function; the independent variable of the optimization objective function is the number of the first candidate backup storage module up to the [number missing]. H The numbers of the candidate backup storage modules; wherein, the first candidate backup storage module to the second... H The value ranges for the candidate backup storage modules are from the first set to the second set. H gather; With the optimization objective function value minimum as the goal, a particle swarm optimization algorithm is used to calculate the first backup storage module number to the H Optimal value of the backup storage module number, respectively, the first optimal number to the H Optimal number; respectively, the first optimal number to the H Optimal number corresponding to the storage module is recorded as the first global optimal backup storage module to the H Global optimal backup storage module.
4. The data security protection method for a civil aviation data hub according to claim 3, wherein, The optimization objective function is: ; in, This represents the objective function to be optimized. Indicates the first The number of the backup storage module to be selected. The value is 1 to Integers between [a certain number] For the first Storage modules to be destroyed and the first Packet loss rate of network links between storage modules To determine the data security level and data storage period based on the pre-set mapping relationship, the first... After the data security level corresponding to the data sub-blocks stored in the storage module is converted into the data storage cycle, it is then processed through a pre-set reverse mapping relationship and normalized to a value within the range of 0 to 1. For the pre-set first weight, The second weight is set in advance; and The sum of is 1.
5. A data security protection system of a civil aviation data center, configured to perform the method of any one of claims 1-4, characterized in that, The system comprises, which are connected in sequence: a first data acquisition module, a data confusion module, a correlation blocking module, and a formatting module; The first data acquisition module is used for obtaining first data; the first data comprises civil aviation multi-source data and correlation characteristic data between the civil aviation multi-source data; the civil aviation multi-source data is divided into high-security-level data and low-security-level data; the correlation characteristic data between the high-security-level data and the low-security-level data in the correlation characteristic data between the civil aviation multi-source data is recorded as first correlation data; The data confusion module is used for dividing the high-security level data into a plurality of data sub-blocks according to data types, and storing the data sub-blocks in a plurality of storage modules. N N N The data confusion module is used for dividing the high-security level data into a plurality of data sub-blocks according to data types, and storing the data sub-blocks in a plurality of storage modules. superimposing the confusion variable with first correlation data to obtain second correlation data; decomposing the confusion variable into N confusion sub-blocks; storing the confusion sub-blocks in N storage modules in a scattered manner; destroying the confusion variable; and opening the storage modules, the second correlation data, and the low-security-level data to user access N N The correlation blocking module is used for generating a restricted access strategy by a correlation blocking method according to the second correlation data and the low-security-level data; the correlation blocking method triggers an abnormal alarm for a request of simultaneously accessing the low-security-level data and the second correlation data; The formatting module is used for, after receiving a formatting instruction, backing up the confused sub-block stored in the storage module receiving the formatting instruction to the storage module not receiving the formatting instruction, and then formatting; The correlation characteristic data between the civil aviation multi-source data is obtained by data mining on the civil aviation multi-source data in advance, and represents the correlation between the civil aviation multi-source data; The method for dividing the civil aviation multi-source data into high-security-level data and low-security-level data comprises: According to a data security level of the civil aviation multi-source data obtained in advance, the civil aviation multi-source data is divided into first security level data to nth security level data Q security level data; Q represents the number of data security levels The first Security level data up to the Q Classified data is designated as high-class data; data classified as first-class data is then classified as high-class data. Data with a security classification level is recorded as low-security data; among which This indicates a pre-set confidentiality threshold; The N storage modules are connected through a distributed network; is a positive integer greater than 2 set in advance. The method for superimposing the confusion variable and the first correlation data to obtain the second correlation data comprises: The superposition formula is used to calculate the second correlation data according to the confusion variable and the first correlation data; and the superposition formula is: ; wherein, represents the first association data, represents the second association data, represents a confounding variable; the method of decomposing the obfuscated variable into N a plurality of obfuscated sub-blocks includes: Adopting a hash function method to generate one independent random number, denoted as a first independent random number to the independent random number ; the value range of the independent random number is between and ; wherein represents a pre-set error control constant with a value greater than 0. According to the above The related disturbance is calculated according to the above ; The calculation formula is: ; wherein, represents the independent random number; is a positive integer having a value of 1 to between 1 and 100. in succession to , cloned as to ; the to are denoted as first offset to nth offset, respectively; According to the confusion variable, the first offset to the offset calculation N confusion sub-blocks, respectively, the first confusion sub-block to the N confusion sub-block; wherein the confusion sub-block calculation formula is: ; wherein, represents the confusion sub-block; is a positive integer having a value of 1 to between 1 and 10.
Citation Information
Patent Citations
Data privacy protection method
CN119848936A
Data security processing method and system based on distributed storage
CN120654250A
Computer application backup method and system
US20040107199A1