Information security management and monitoring system based on big data
Through the information security management system based on big data, real-time monitoring of equipment performance, automatic analysis and generation of desensitization strategies, efficient sensitive data processing and multiple encryption are achieved, solving the problems of low efficiency and insufficient security strategies in traditional desensitization processing, and improving the overall efficiency and security of information security management.
Patent Information
- Application Number
- CN202510530624.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional sensitive data desensitization processing is inefficient and difficult to accurately identify risks, resulting in high data processing costs and increased redundant data. The existing security strategies lack multiple encryption and dynamic protection, which cannot meet modern information security needs.
It adopts information security management and monitoring systems based on big data, including equipment monitoring modules, data acquisition modules, data desensitization modules, library management modules and security prevention and control modules, monitors equipment performance in real time, automatically collects and analyzes sensitive data, generates desensitization strategies based on security levels, and ensures data security through multiple encryption and distributed storage.
Real-time risk warning is realized, data processing efficiency is improved, manual intervention costs are reduced, storage resource use is optimized, data multiple encryption protection capabilities are enhanced, and modern information security needs are met.
Smart Images

Figure CN120408711A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly to an information security management and monitoring system based on big data. Background Art
[0002] In the digital age, with the rapid development of big data, the status of information security management has become increasingly important. Big data information security management is not only a necessary means to protect sensitive data, but also the cornerstone to ensure enterprise compliance, maintain user trust, and achieve sustainable development.
[0003] Desensitization refers to the process of processing sensitive data to remove or hide the personally identifiable information therein, so as to protect data privacy and security. In traditional technologies, the desensitization processing of sensitive data often adopts static rules or fixed strategies. This method is not only inefficient, but also prone to waste of data processing costs. Especially when facing texts with insufficient sensitivity, traditional technologies are difficult to accurately identify and evaluate their actual risks, resulting in poor data desensitization effects. Due to the lack of effective classification and grading, enterprises often need to invest a large amount of resources in manual review and modification when implementing desensitization strategies, increasing the management cost.
[0004] Secondly, redundant data may be generated during the desensitization process. These redundant information not only occupies storage resources, but also may increase the complexity of subsequent data analysis, making it more difficult to extract valuable information. Moreover, the existing security strategies lack the support of multiple encryption and dynamic protection, making sensitive data extremely vulnerable to attacks during transmission and storage, and unable to meet the requirements of modern information security. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an information security management and monitoring system based on big data to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: An information security management and monitoring system based on big data, comprising:
[0007] A device monitoring module, which is used to monitor the hardware devices storing sensitive data in real time to obtain an operation data set, and construct a device performance fluctuation coefficient Wx. If the device performance fluctuation coefficient Wx exceeds the fluctuation threshold Q, a risk warning is initially sent outwards;
[0008] The big data collection module automatically collects and scans several data sources, and after uniformly converting each data source into an editable first text, extracts the sensitive features in the first text and the interaction degree of the collected sensitive words on social media, and analyzes and calculates to obtain: the number of sensitive words Mgcsl appearing per hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the public exposure Gkd of the sensitive words on social media. After correlation, a text sensitivity index MGbh is constructed and evaluated to obtain the corresponding first security level, second security level, and third security level;
[0009] The data desensitization module is used to generate and execute corresponding desensitization strategies according to the first security level, second security level, and third security level to form a desensitized data set;
[0010] The sub-library management module is used to divide the desensitized data set into several sub-libraries, distribute the several sub-libraries to different physical nodes, and then encrypt the desensitized data set to generate the first encrypted data, and construct a sensitive data protection coefficient Kp;
[0011] The security prevention and control module is used to preset a security threshold Pc, and compare and evaluate the sensitive data protection coefficient Kp with the security threshold Pc. If the sensitive data protection coefficient Kp is lower than the security threshold Pc, the first encrypted data is processed twice to obtain the second encrypted data.
[0012] Preferably, the device monitoring module includes a real-time monitoring unit, a first calculation unit, and a first comparison unit;
[0013] The real-time monitoring unit is used to collect and obtain the running status of the hardware device storing sensitive data in real time through a performance analysis tool Perf, Java, or VisualVM to obtain a running data set, and the running data set includes the thread increment Xczl, the thread blocking time Xczs, the concurrent connection number Bfljs, the cache hit rate Hcmz, the context switching rate Sxqh, and the CPU load Fzl;
[0014] The first calculation unit is used to preprocess and dimensionless process the running data set, and calculate and obtain the device performance fluctuation coefficient Wx through the following formula:
[0015]
[0016] In the formula, represents the maximum thread increment security threshold, represents the preset maximum thread blocking time security threshold, represents the maximum security threshold for the concurrent connection number, represents the maximum security threshold for the context switching rate, It represents the maximum safe threshold of CPU load, and w1, w2, w3, w4, and w5 represent weight values;
[0017] The first comparison unit is used to preset a fluctuation threshold Q, and compare and evaluate the device performance fluctuation coefficient Wx with the fluctuation threshold Q to determine whether there is an attack risk in the operating state of the hardware device storing sensitive data, including:
[0018] If the device performance fluctuation coefficient Wx > the fluctuation threshold Q, it indicates that there is an attack risk in the hardware device storing sensitive data, and a risk warning is initially sent outwards;
[0019] If the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q, it indicates that there is no attack risk in the hardware device storing sensitive data. The devices with the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q are classified into the safe device group, and information collection is carried out based on big data using the safe device group.
[0020] Preferably, the big data collection module includes a data source collection unit and an identification unit;
[0021] The data source collection unit is used to automatically collect and scan several data sources through network protocols, APIs, or file systems using the safe device group, and identify and classify the several data sources obtained through collection through the identification unit;
[0022] The steps of identifying and classifying several data sources include:
[0023] S11. Analyze the formats of the several scanned data sources to identify the data types. The data types include JSON, XML, CSV, text, and images; and uniformly convert different data types into editable first texts;
[0024] S12. Generate metadata for each editable first text, including data source name, metadata format, data size, creation time, update frequency, and classification information.
[0025] Preferably, the identification unit includes a sensitive word extraction unit and a second calculation unit;
[0026] The sensitive word extraction unit is used to extract sensitive features of the data content for each editable first text. The sensitive features of the data content include personal identity information, contact information, financial information, medical information, business secret information, and legal information;
[0027] Deeply analyze the sensitive features of the data content to obtain the number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of the sensitive word sentences;
[0028] The number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of the sensitive word sentences are calculated and obtained through the following formulas:
[0029]
[0030]
[0031] In the formula, C s is the total number of sensitive words in the text, T w is the total number of words in the text, C i is the number of occurrences of the i-th type of sensitive word in the text, n is the total number of sensitive word types, L s is the total number of words in the sentences containing sensitive words, S s is the total number of sentences containing sensitive words;
[0032] The second calculation unit is used to collect the interaction degree of the content containing each sensitive word on the public social media, including the number of likes, shares, and comments, so as to construct the publicity Gkd of the sensitive word on the social media:
[0033]
[0034] In the formula, N t represents the number of posts containing sensitive words, N s represents all the collected relevant posts, including the number of posts without sensitive words, and T represents the total interaction number obtained by the content containing sensitive words. The total interaction number includes the sum of the number of likes, shares, and comments.
[0035] Preferably, the recognition unit further includes an association unit and a second evaluation unit;
[0036] The association unit is used to extract the number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive word on the social media. After dimensionless processing, the text sensitivity index MGbh is calculated and obtained through the following association formula:
[0037]
[0038] In the formula, r1, r2, r3, and r4 respectively represent the weight values of the number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive word on the social media, and 0 < r1 < 1, 0 < r2 < 1, 0 < r3 < 1, 0 < r4 < 1. Their specific values are adjusted and set by the user, and r1 + r2 + r3 + r4 = 1; E represents a correction constant.
[0039] Preferably, the second evaluation unit is configured to preset a first security level threshold X1 and a second security level threshold X2, where the first security level threshold X1 > the second security level threshold X2, and compare and evaluate the text sensitivity index MGbh with the first security level threshold X1 and the second security level threshold X2 respectively to determine the security level corresponding to the text sensitivity index MGbh of the first text, including:
[0040] If the text sensitivity index MGbh < the second security level threshold X2, it indicates that the sensitivity of the first text is low, and a first security level is generated;
[0041] If the second security level threshold X2 ≤ the text sensitivity index MGbh ≤ the first security level threshold X1, it indicates that the sensitivity of the first text is medium, and a second security level is generated;
[0042] If the text sensitivity index MGbhi > the first security level threshold X1; it indicates that the sensitivity of the first text is high, and a third security level is generated.
[0043] Preferably, the data desensitization module includes a feature encoding unit and a desensitization unit;
[0044] The feature encoding unit is configured to generate corresponding desensitization strategies according to the first security level, the second security level, and the third security level and execute them, including:
[0045] Generate a first desensitization strategy according to the first security level, including: converting 50% of the sensitive fields in the first text into an anonymous coding format, specifically replacing 50% of the sensitive fields in the first text with the symbol "#" and retaining 50% of the original data;
[0046] Generate a second desensitization strategy according to the second security level, including: converting 30% of the sensitive fields in the first text into an anonymous coding format, specifically randomly generated fictional data, and covering 70% of the sensitive fields with black blocks;
[0047] Generate a third desensitization strategy according to the third security level, including: performing the most stringent desensitization process, covering 90% of the sensitive fields in the first text, converting the remaining 10% of the sensitive fields into an anonymous coding format, specifically replacing them with the character "@", and performing a hash process on the first text to generate a first key, and the first text can be read only after obtaining the resolution of the first key;
[0048] The desensitization unit is configured to count the first text converted by the feature encoding unit, convert the sensitive fields into non-sensitive fields to form a second text, and establish a desensitized data set.
[0049] Preferably, the sub-library management module includes a data partitioning unit, a distributed storage unit, an encryption processing unit, and a data analysis unit;
[0050] The data partitioning unit is used to partition the desensitized data set into a first word library, a second word library, and a third word library according to the first security level, the second security level, and the third security level;
[0051] The distributed storage unit stores the first word library, the second word library, and the third word library on different physical nodes;
[0052] The encryption processing unit uses a combined algorithm of AES and RSA to encrypt the desensitized data sets of the first word library, the second word library, and the third word library, generates first encrypted data, and stores it in the supervision database;
[0053] The data analysis unit is used to extract the valid value features and redundant value features of the desensitized and encrypted desensitized data set, collect and obtain the user access frequency Fwpl for statistical analysis, and calculate and obtain the sensitive data protection coefficient Kp through the following formula:
[0054]
[0055] In the formula, V t represents the valid value of the desensitized data, V r represents the redundant value of the desensitized data, C represents the total number of desensitized information included, α represents the sensitive information weight coefficient, R represents the number of sensitive information that can be retrieved when restoring the data, and β represents the information recovery ability influence coefficient.
[0056] Preferably, the security prevention and control module is used to preset a security threshold Pc, compare and evaluate the sensitive data protection coefficient Kp with the security threshold Pc, and generate a security evaluation result, including:
[0057] If the sensitive data protection coefficient Kp > the security threshold Pc, it means that the protection ability for the desensitized data set is qualified;
[0058] If the sensitive data protection coefficient Kp ≤ the security threshold Pc, it means that the protection ability for the desensitized data set is unqualified, and a protection strategy is generated, including: deleting the redundant value V r of the desensitized data, and performing secondary processing on the first encrypted data. The secondary processing includes splitting processing and additional encryption processing. The splitting processing includes splitting the first encrypted data into multiple data blocks, and applying an independent second key to each data block. The second key is dynamically rotated after 10 - 15 user accesses;
[0059] The additional encryption process includes: applying an additional encryption algorithm, such as Blowfish or ChaCha20, to the first encrypted data, and adding a hash checksum and a digital signature during the encryption process to generate the second encrypted data.
[0060] Preferably, the first encrypted data includes: user identity information, access log information, and data encryption meta-information;
[0061] The data encryption meta-information includes the encryption algorithm type, key version, and encryption time;
[0062] The second encrypted data includes: the result information after applying Blowfish or ChaCha20 to the first encrypted data, independent key-related information generated for each data block, and the hash value and digital signature.
[0063] The present invention provides an information security management and monitoring system based on big data. It has the following beneficial effects:
[0064] (1) Real-time monitoring of hardware devices storing sensitive data, quickly obtaining the operation data set, and promptly constructing the device performance fluctuation coefficient Wx. If performance anomalies are detected, the system can immediately issue a risk warning, thus quickly responding to potential attacks. This real-time nature effectively reduces the response time to security risks and enhances the overall security of the system.
[0065] (2) The big data collection module automatically collects and uniformly converts multiple data sources into editable first texts, improving data processing efficiency. By extracting sensitive features and the degree of social media interaction, this module can quantify the influence of sensitive information and construct the text sensitivity index MGbh based on this. This analysis not only optimizes the identification of sensitive data but also provides data support for subsequent desensitization strategies, reducing the cost of manual intervention.
[0066] (3) The data desensitization module generates corresponding desensitization strategies according to different security levels and automatically processes sensitive data. This process reduces manual review, improves desensitization efficiency, ensures privacy protection during data use, and thus reduces the legal and economic risks that may be caused by data leakage.
[0067] (4) The sub-library management module divides the desensitized data set into multiple sub-libraries and distributes them for storage on different physical nodes, which not only reduces the risk of a single library but also optimizes the use of storage resources and reduces the generation of redundant data. This mechanism enhances the elasticity and scalability of the system. The security prevention and control module can automatically identify situations with insufficient protection capabilities by presetting a security threshold Pc and comparing it with the sensitive data protection coefficient Kp, and take secondary processing measures. This mechanism enhances the multiple encryption protection ability of the data and ensures the security of sensitive data during storage and transmission, meeting the requirements of modern information security. Brief Description of the Drawings
[0068] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Embodiments
[0069] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] Embodiment 1
[0071] Please refer to Figure 1 , the present invention provides an information security management and monitoring system based on big data, including:
[0072] An equipment monitoring module for real-time monitoring of hardware devices storing sensitive data to obtain an operation data set and construct an equipment performance fluctuation coefficient Wx. If the equipment performance fluctuation coefficient Wx exceeds the fluctuation threshold Q, a risk warning is preliminarily sent outwards;
[0073] A big data collection module automatically collects and scans a number of data sources, and after uniformly converting each data source into an editable first text, extracts the sensitive features in the first text and the interaction degree of the collected sensitive words on social media, and analyzes and calculates to obtain: the number of sensitive words per hundred words Mgcsl, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive words on social media. After correlation, a text sensitivity index MGbh is constructed and evaluated to obtain the corresponding first security level, second security level, and third security level;
[0074] A data de-sensitization module for generating and executing corresponding de-sensitization strategies according to the first security level, second security level, and third security level to form a de-sensitized data set;
[0075] A sub-library management module for dividing the de-sensitized data set into several sub-libraries, distributing the several sub-libraries in different physical nodes, encrypting the de-sensitized data set to generate first encrypted data, and constructing a sensitive data protection coefficient Kp;
[0076] A security prevention and control module for presetting a security threshold Pc and comparing and evaluating the sensitive data protection coefficient Kp with the security threshold Pc. If the sensitive data protection coefficient Kp is lower than the security threshold Pc, the first encrypted data is processed twice to obtain second encrypted data.
[0077] In this embodiment, the hardware device storing sensitive data is monitored in real time, the operation data set is quickly obtained, and the device performance fluctuation coefficient Wx is constructed in a timely manner. If a performance anomaly is detected, the system can immediately issue a risk warning, thus quickly responding to potential attacks. This real-time feature effectively reduces the response time to security risks and improves the overall security of the system.
[0078] The big data collection module automatically collects and uniformly converts multiple data sources into editable first texts, improving data processing efficiency. By extracting sensitive features and the degree of social media interaction, this module can quantify the influence of sensitive information and construct the text sensitivity index MGbh based on this. This analysis not only optimizes the identification of sensitive data but also provides data support for subsequent desensitization strategies, reducing the cost of manual intervention.
[0079] The data desensitization module generates corresponding desensitization strategies according to different security levels and automatically processes sensitive data. This process reduces manual review, improves desensitization efficiency, ensures privacy protection during data use, and thus reduces the legal and economic risks that may be caused by data leakage.
[0080] The sub-library management module divides the desensitized data set into multiple sub-libraries and distributes them for storage on different physical nodes, which not only reduces the risk of a single library but also optimizes the use of storage resources and reduces the generation of redundant data. This mechanism enhances the elasticity and scalability of the system. The security prevention and control module can automatically identify situations with insufficient protection capabilities by presetting the security threshold Pc and comparing and evaluating it with the sensitive data protection coefficient Kp, and take secondary processing measures. This mechanism enhances the multi-layer encryption protection ability of the data, ensures the security of sensitive data during storage and transmission, and meets the requirements of modern information security.
[0081] Embodiment 2
[0082] This embodiment is an explanatory description based on Embodiment 1. Please refer to Figure 1 , specifically, the device monitoring module includes a real-time monitoring unit, a first calculation unit, and a first comparison unit;
[0083] The real-time monitoring unit is used to collect and obtain the operation status of the hardware device storing sensitive data in real time through the performance analysis tools Perf, Java, or VisualVM to obtain the operation data set, and the operation data set includes the thread increment Xczl, the thread blocking time Xczs, the concurrent connection number Bfljs, the cache hit rate Hcmz, the context switching rate Sxqh, and the CPU load Fzl; This real-time monitoring improves the transparency of system performance, enabling potential performance bottlenecks and anomalies to be discovered in a timely manner, which helps to reduce system failures and improve resource utilization.
[0084] The first calculation unit is used to preprocess and dimensionless process the operation data set, and calculate and obtain the device performance fluctuation coefficient Wx through the following formula:
[0085]
[0086] In the formula, represents the maximum thread increment safety threshold, represents the preset maximum thread blocking time safety threshold, represents the maximum safety threshold of the concurrent connection number, represents the maximum safety threshold of the context switching rate, represents the maximum safety threshold of the CPU load, and w1, w2, w3, w4, and w5 represent weight values; through the above calculation method, the performance state of the device can be quantified, which is convenient for horizontal comparison and evaluation. This process reduces the necessity of human intervention, improves the accuracy and consistency of data processing, and provides a solid foundation for subsequent risk assessment.
[0087] The first comparison unit is used to preset a fluctuation threshold Q, and compare and evaluate the device performance fluctuation coefficient Wx with the fluctuation threshold Q to determine whether there is an attack risk in the operation state of the hardware device storing sensitive data, including:
[0088] If the device performance fluctuation coefficient Wx > the fluctuation threshold Q, it means that there is an attack risk in the hardware device storing sensitive data. A risk warning is initially sent outwards, and the access permission to sensitive data of the hardware device is temporarily restricted to prevent possible leakage of sensitive information, and the suspicious device or service is isolated to ensure that potential attacks will not affect other systems or data;
[0089] If the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q, it means that there is no attack risk in the hardware device storing sensitive data. The devices with the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q are classified into the safe device group, and information collection is carried out based on big data by the safe device group.
[0090] In this embodiment, by presetting a fluctuation threshold Q and comparing it with the device performance fluctuation coefficient Wx, it is possible to effectively determine whether there is an attack risk for the hardware device storing sensitive data. This mechanism not only improves the accuracy of risk identification but also provides a basis for taking corresponding security measures in a timely manner to ensure the security of sensitive data. If it is found that the device performance fluctuation coefficient Wx exceeds the preset fluctuation threshold Q, the system will immediately send out a risk warning and temporarily restrict the access permission of the hardware device to sensitive data. This mechanism can effectively prevent potential leakage of sensitive information, protect the integrity and confidentiality of data, and at the same time isolate suspicious devices or services to reduce the impact on other systems or data. Devices with the performance fluctuation coefficient Wx within the security threshold are classified into a safe device group, and a big data-based information collection method is adopted. This classification method improves the management efficiency of the system, helps to better utilize resources, and realizes centralized monitoring and management of safe devices, enhancing the overall security of the system.
[0091] Embodiment 3
[0092] This embodiment is an explanatory description carried out in Embodiment 1. Please refer to Figure 1 , specifically, the big data collection module includes a data source collection unit and an identification unit;
[0093] The data source collection unit is used to automatically collect and scan a number of data sources through network protocols, APIs or file systems using the safe device group, and identify and classify the number of data sources obtained through collection through the identification unit;. This automated process reduces the need for manual intervention, reduces the time cost of data collection, and at the same time ensures the diversity and integrity of data sources, laying a solid foundation for subsequent data analysis.
[0094] The steps of identifying and classifying a number of data sources include:
[0095] S11. Analyze the formats of the scanned number of data sources to identify data types, and the data types include JSON, XML, CSV, text, and images; and uniformly convert different data types into editable first texts; this process not only improves the standardization of data processing but also provides convenient conditions for subsequent sensitive feature extraction and data analysis, helping to improve the efficiency and quality of data processing.
[0096] S12. Generate metadata for each editable first text, including data source name, metadata format, data size, creation time, update frequency, and classification information. The generation of this metadata makes the management and retrieval of data more efficient, facilitates subsequent data analysis and tracking, and enhances the availability and operability of data.
[0097] The identification unit includes a sensitive word extraction unit and a second calculation unit;
[0098] The sensitive word extraction unit is used to extract sensitive features of data content for each editable first text. The sensitive features of data content include personal identity information, contact information, financial information, medical information, business secret information, and legal information. The system can comprehensively identify and classify sensitive data. This comprehensive extraction of sensitive features ensures the accuracy of data processing, can effectively reduce the risk of data leakage, and provides basic support for subsequent desensitization strategies.
[0099] Perform in-depth analysis on the sensitive features of data content to obtain the number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of sensitive word sentences. This quantitative analysis provides an intuitive data basis for enterprises to evaluate the risk level of texts, helps enterprises prioritize the processing of high-risk content, and thus improves the efficiency of data security management.
[0100] The number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of sensitive word sentences are obtained through the following formulas:
[0101]
[0102] In the formula, C s is the total number of sensitive words in the text, T w is the total number of words in the text, C i is the number of occurrences of the i-th type of sensitive word in the text, n is the total number of sensitive word types, L s is the total number of words in the sentences containing sensitive words, S s is the total number of sentences containing sensitive words;
[0103] The second calculation unit is used to collect the interaction degree of each sensitive word content on public social media, including the number of likes, shares, and comments, to construct the publicity degree Gkd of sensitive words on social media:
[0104]
[0105] In the formula, N t represents the number of posts containing sensitive words, N s represents all the collected relevant posts, including the number of posts without sensitive words, and T represents the total interaction number obtained by the content containing sensitive words. The total interaction number includes the sum of the number of likes, shares, and comments. This analysis can evaluate the spread degree of sensitive content among the public, help enterprises understand the external environment of their data risks, facilitate the adoption of corresponding protection measures, and enhance the security of data.
[0106] The recognition unit further includes an association unit and a second evaluation unit;
[0107] The association unit is used to extract the number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive words on social media. After dimensionless processing, the text sensitivity index MGbh is calculated through the following association formula:
[0108]
[0109] In the formula, r1, r2, r3, and r4 respectively represent the weight values of the number of sensitive words Mgcsl appearing in every hundred words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive words on social media, and 0 < r1 < 1, 0 < r2 < 1, 0 < r3 < 1, 0 < r4 < 1. Their specific values are adjusted and set by the user, and r1 + r2 + r3 + r4 = 1; E represents a correction constant. The text sensitivity index MGbh provides an enterprise with a quantitative way to evaluate the sensitivity of the text, making the identification and management of sensitive data more scientific and efficient. Allowing users to adjust the weight values (r1, r2, r3, r4) according to specific needs makes the system more flexible and able to adapt to different scenarios and risk assessment requirements. This flexibility enhances the applicability of the system, and enterprises can optimize sensitivity analysis according to the actual business environment and data characteristics, improving the pertinence of data processing. Through the calculation of the text sensitivity index, enterprises can quickly identify high-risk texts and formulate corresponding desensitization strategies or security measures accordingly. This index not only provides a quantitative risk assessment tool but also provides strong support for the decision-making process of enterprises, ensuring that information security management is more forward-looking and effective.
[0110] The second evaluation unit is used to preset a first security level threshold X1 and a second security level threshold X2, and the first security level threshold X1 > the second security level threshold X2, and compare and evaluate the text sensitivity index MGbh with the first security level threshold X1 and the second security level threshold X2 respectively to determine the security level corresponding to the text sensitivity index MGbh of the first text, including:
[0111] If the text sensitivity index MGbh < the second security level threshold X2, it indicates that the sensitivity of the first text is low, and a first security level is generated;
[0112] If the second security level threshold X2 ≤ the text sensitivity index MGbh ≤ the first security level threshold X1, it indicates that the sensitivity of the first text is medium, and a second security level is generated;
[0113] If the text sensitivity index MGbhi > the first security level threshold X1, it indicates that the first text has a high sensitivity, and the third security level is generated.
[0114] In this embodiment, through the preset security level thresholds (X1 and X2), the second evaluation unit can automatically classify the text sensitivity index (MGbh). This function simplifies the processing flow of sensitive data, enabling enterprises to quickly identify and process data with different sensitivity levels, thereby improving work efficiency. The system can accurately judge the sensitivity of the text. According to different ranges of MGbh, corresponding security levels (the first, second, or third security level) are generated, providing enterprises with accurate risk assessment and avoiding misjudgments that may occur in traditional methods. By classifying the text sensitivity, enterprises can allocate resources reasonably. For example, for highly sensitive texts, strict desensitization processing and security monitoring are prioritized, while for low-sensitivity texts, the monitoring frequency and processing intensity can be reduced. This optimization helps improve resource utilization efficiency and reduce unnecessary management costs.
[0115] Embodiment 4
[0116] This embodiment is an explanatory description based on Embodiment 3. Please refer to Figure 1 , specifically, the data desensitization module includes a feature encoding unit and a desensitization unit;
[0117] The feature encoding unit is used to generate corresponding desensitization strategies according to the first security level, the second security level, and the third security level and execute them, including:
[0118] Generating a first desensitization strategy according to the first security level, including: converting 50% of the sensitive fields in the first text into an anonymous coding format, specifically replacing 50% of the sensitive fields in the first text with the symbol "#" and retaining 50% of the original data;
[0119] Generating a second desensitization strategy according to the second security level, including: converting 30% of the sensitive fields in the first text into an anonymous coding format, specifically randomly generated fictional data, and covering 70% of the sensitive fields with black blocks;
[0120] Generating a third desensitization strategy according to the third security level, including: performing the strictest desensitization processing, covering 90% of the sensitive fields in the first text, converting the remaining 10% of the sensitive fields into an anonymous coding format, specifically replacing them with the character "@", and performing a hash processing on the first text to generate a first key, and the first text can be read only after obtaining the resolution of the first key;
[0121] The desensitization unit is used to count the first text converted by the feature encoding unit, convert the sensitive fields into non-sensitive fields to form a second text, and establish a desensitized data set.
[0122] In this embodiment, the feature encoding unit generates corresponding desensitization strategies according to different security levels to ensure a flexible processing method based on data sensitivity. This hierarchical strategy can effectively balance data protection and information availability, improving the adaptability of the desensitization process. The strict desensitization process for the third security level (such as 90% sensitive field occlusion and hashing) greatly improves the security of the data. By deeply processing sensitive information, the risk of information leakage is reduced, ensuring that enterprises can effectively protect important data when facing potential attacks. At the first and second security levels, part of the original data is retained and anonymous coding processing is adopted, enabling effective data analysis and mining without compromising data privacy. This provides the possibility for enterprises to utilize data while protecting user privacy. The desensitization unit can automatically count and convert sensitive fields into non-sensitive fields to form a desensitized data set. This automated process reduces manual intervention, improves work efficiency, and reduces the risk of data processing errors caused by human mistakes. By implementing effective desensitization measures, enterprises can better comply with data protection regulations and enhance compliance. In addition, transparent and secure data processing enhances customers' trust in the enterprise, making them more willing to share sensitive information with the enterprise.
[0123] Embodiment 5
[0124] This embodiment is an explanatory description based on Embodiment 1. Please refer to Figure 1 , specifically, the sub-database management module includes a data partitioning unit, a distributed storage unit, an encryption processing unit, and a data analysis unit;
[0125] The data partitioning unit is used to partition the desensitized data set into a first sub-database, a second sub-database, and a third sub-database according to the first security level, the second security level, and the third security level to reduce the risk of a single database; the data partitioning unit partitions the desensitized data set into multiple sub-databases according to the security level, effectively reducing the risk when a single data storage location is attacked. This sub-database strategy improves data security and ensures that losses are controllable in the event of a security incident.
[0126] The distributed storage unit stores the first sub-database, the second sub-database, and the third sub-database on different physical nodes; by storing the sub-databases on different physical nodes through the distributed storage unit, data redundancy can be effectively managed, and the availability and fault tolerance of the data can be improved. This architecture design enables other nodes to still provide data services when a certain node fails.
[0127] The encryption processing unit uses a combined algorithm of AES and RSA to encrypt the desensitized data sets of the first sub-library, the second word library, and the third word library, generating the first encrypted data and storing it in the supervision database; the encryption processing unit uses a combined algorithm of AES and RSA to strongly encrypt the desensitized data to ensure the security of sensitive information during storage. This multi-level encryption method improves the data protection ability and effectively blocks unauthorized access.
[0128] The data analysis unit is used to extract the valid value features and redundant value features of the desensitized and encrypted desensitized data sets, collect and obtain the user access frequency Fwpl for statistical analysis, and calculate and obtain the sensitive data protection coefficient Kp through the following formula:
[0129]
[0130] In the formula, V t represents the valid value of the desensitized data, V r represents the redundant value of the desensitized data, C represents the total number of desensitized information included, α represents the sensitive information weight coefficient, R represents the number of sensitive information that can be retrieved when restoring the data, and β represents the information recovery ability influence coefficient.
[0131] In this embodiment, the data analysis unit can comprehensively evaluate the data usage situation by extracting the valid value features and redundant value features and statistically analyzing the user access frequency. The analysis of the sensitive data protection coefficient Kp helps the enterprise understand the actual utilization status of the data, optimize the data management strategy, and improve the data usage efficiency. By calculating the sensitive data protection coefficient Kp, the security of the desensitized data set can be dynamically evaluated. This indicator provides feedback for the enterprise on the data protection measures, enabling it to adjust and optimize the security strategy in real time to ensure the continuity and effectiveness of information security.
[0132] Embodiment 6
[0133] This embodiment is an explanatory description based on Embodiment 5. Please refer to Figure 1 , specifically, the security prevention and control module is used to preset a security threshold Pc, compare and evaluate the sensitive data protection coefficient Kp with the security threshold Pc, and generate a security evaluation result, including:
[0134] If the sensitive data protection coefficient Kp > the security threshold Pc, it indicates that the protection ability for the desensitized data set is qualified;
[0135] If the sensitive data protection coefficient Kp ≤ the security threshold Pc, it indicates that the protection ability for the desensitized data set is unqualified, and a protection strategy is generated, including: deleting the redundant value V of the desensitized data r, and perform secondary processing on the first encrypted data. The secondary processing includes segmentation processing and additional encryption processing. The segmentation processing includes splitting the first encrypted data into multiple data blocks and applying an independent second key to each data block. The second key is dynamically rotated after 10 - 15 user accesses;
[0136] The additional encryption processing includes: applying an additional encryption algorithm, Blowfish or ChaCha20, to the first encrypted data. During the encryption process, a hash checksum and a digital signature are added to generate the second encrypted data.
[0137] Specifically, the first encrypted data includes: user identity information, access log information, and data encryption meta - information;
[0138] The data encryption meta - information includes the encryption algorithm type, key version, and encryption time;
[0139] The second encrypted data includes: the result information after applying Blowfish or ChaCha20 to the first encrypted data, the independent key - related information generated for each data block, and the hash value and digital signature.
[0140] In this embodiment, the security prevention and control module evaluates the protection ability of the de - sensitized data set in real - time by comparing the sensitive data protection coefficient Kp with the security threshold Pc. This dynamic evaluation mechanism ensures the timely monitoring of the data security status and helps enterprises quickly identify potential risks. When the protection ability is unqualified, the module can automatically generate corresponding protection strategies, including deleting redundant values and performing secondary processing on the data. This automated processing reduces the need for human intervention and improves the response speed and management efficiency. By splitting the first encrypted data into multiple data blocks through segmentation processing and applying an independent second key to each data block, the security of the data can be significantly enhanced. Even if a certain data block is attacked, the attacker cannot obtain the information of the entire data set, reducing the risk of information leakage. The second key is dynamically rotated after the user access frequency reaches a specific threshold, improving the flexibility and security of key management. This mechanism effectively prevents potential risks caused by the long - term exposure of the key. By applying additional encryption algorithms such as Blowfish or ChaCha20 and combining hash checksums and digital signatures, the generated second encrypted data is further enhanced in security. This multiple protection strategy ensures the security of sensitive data during storage and transmission, preventing unauthorized access and data tampering. The management of data encryption meta - information, including the encryption algorithm type, key version, and encryption time, provides an important basis for subsequent data recovery and auditing. This detailed meta - information record helps enhance the traceability and security compliance of the data.
[0141] The setting of the threshold value is for the convenience of comparison. Regarding the size of the threshold value, it depends on the amount of sample data and the base quantity set by those skilled in the art for each group of sample data; as long as the proportional relationship between the parameter and the quantized value is not affected.
[0142] The above formulas are all obtained by collecting a large amount of data for software simulation and selecting a formula close to the true value. The coefficients in the formula are set by those skilled in the art according to the actual situation. As mentioned above, the above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An information security management and monitoring system based on big data, characterized in that, Including: The device monitoring module is used to monitor the hardware device storing sensitive data in real time to obtain an operation data set and construct a device performance fluctuation coefficient Wx. If the device performance fluctuation coefficient Wx exceeds the fluctuation threshold Q, a risk warning is preliminarily sent outwards. The big data collection module automatically collects and scans several data sources. After uniformly converting each data source into an editable first text, it extracts the sensitive features in the first text and the interaction degree of the collected sensitive words on social media, and analyzes and calculates to obtain: the number of sensitive words per hundred words Mgcsl, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of the sensitive word sentences, and the publicity Gkd of the sensitive words on social media. After correlation, a text sensitivity index MGbh is constructed and evaluated to obtain the corresponding first security level, second security level, and third security level. The data desensitization module is used to generate and execute corresponding desensitization strategies according to the first security level, second security level, and third security level to form a desensitized data set. The sub-library management module is used to divide the desensitized data set into several sub-libraries, distribute the several sub-libraries to different physical nodes, then encrypt the desensitized data set to generate first encrypted data, and construct a sensitive data protection coefficient Kp. The security prevention and control module is used to preset a security threshold Pc, and compare and evaluate the sensitive data protection coefficient Kp with the security threshold Pc. If the sensitive data protection coefficient Kp is lower than the security threshold Pc, the first encrypted data is secondarily processed to obtain second encrypted data.
2. The information security management and monitoring system based on big data according to claim 1, wherein, The device monitoring module includes a real-time monitoring unit, a first calculation unit, and a first comparison unit. The real-time monitoring unit is used to collect and obtain the operation status of the hardware device storing sensitive data in real time through a performance analysis tool Perf, Java, or VisualVM to obtain an operation data set, and the operation data set includes a thread increment Xczl, a thread blocking time Xczs, a concurrent connection number Bfljs, a cache hit rate Hcmz, a context switching rate Sxqh, and a CPU load Fzl. The first calculation unit is used to preprocess and dimensionless process the operation data set, and calculate and obtain the device performance fluctuation coefficient Wx through the following formula: In the formula, represents the maximum thread increment safety threshold, represents the preset maximum thread blocking time safety threshold, represents the maximum safety threshold of the concurrent connection number, represents the maximum safety threshold of the context switching rate, represents the maximum safety threshold of the CPU load, and w1, w2, w3, w4, and w5 represent weight values; The first comparison unit is used to preset a fluctuation threshold Q, and compare and evaluate the device performance fluctuation coefficient Wx with the fluctuation threshold Q to determine whether there is an attack risk in the operation status of the hardware device storing sensitive data, including: If the device performance fluctuation coefficient Wx > the fluctuation threshold Q, it indicates that there is an attack risk in the hardware device storing sensitive data, and a risk warning is preliminarily sent outwards. If the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q, it indicates that there is no attack risk in the hardware device storing sensitive data, and the device with the device performance fluctuation coefficient Wx ≤ the fluctuation threshold Q is classified into a safe device group, and information collection is performed based on big data using the safe device group.
3. An information security management and monitoring system based on big data according to claim 1, characterized in that, The big data collection module includes a data source collection unit and an identification unit. The data source acquisition unit is used to automatically collect and scan several data sources through network protocols, APIs, or file systems using a group of security devices, and identify and classify the several data sources obtained through collection through the identification unit; The steps of identifying and classifying several data sources include: S11. Analyze the formats of the several scanned data sources to identify data types, where the data types include JSON, XML, CSV, text, and images; and uniformly convert different data types into editable first texts; S12. Generate metadata for each editable first text, including data source name, metadata format, data size, creation time, update frequency, and classification information.
4. An information security management and monitoring system based on big data according to claim 3, characterized in that The identification unit includes a sensitive word extraction unit and a second calculation unit; The sensitive word extraction unit is used to extract sensitive features of data content from each editable first text, where the sensitive features of data content include personal identity information, contact information, financial information, medical information, trade secret information, and legal information; Deeply analyze the sensitive features of data content to obtain the number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of sensitive word sentences; The number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, and the average length pjcs of sensitive word sentences are obtained through the following formulas: Wherein, C s is the total number of sensitive words in the text, T w is the total number of words in the text, C i is the number of occurrences of the i-th sensitive word in the text, n is the total number of sensitive word types, L s is the total number of words in the sentences containing sensitive words, S s is the total number of sentences containing sensitive words; The second calculation unit is used to collect the interaction levels of each sensitive word content on public social media, including the number of likes, shares, and comments, to construct the public degree Gkd of sensitive words on social media: Where N t represents the number of posts containing sensitive words, and N s represents the total number of all collected relevant posts, including the number of posts that do not contain sensitive words. T represents the total number of interactions obtained by the content containing sensitive words, and the total number of interactions includes the sum of the number of likes, shares, and comments.
5. An information security management and monitoring system based on big data according to claim 4, characterized in that, The identification unit further includes an association unit and a second evaluation unit; The association unit is used to extract the number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of sensitive word sentences, and the public degree Gkd of sensitive words on social media. After dimensionless processing, calculate the text sensitivity index MGbh through the following association formula: In the formula, r1, r2, r3, and r4 respectively represent the weight values of the number of sensitive words Mgcsl per 100 words, the number of repetitions Cfcs of the same type of sensitive words in a single text, the average length pjcs of sensitive word sentences, and the public degree Gkd of sensitive words on social media, and 0 < r1 < 1, 0 < r2 < 1, 0 < r3 < 1, 0 < r4 < 1. Their specific values are adjusted and set by the user, and r1 + r2 + r3 + r4 = 1; E represents a correction constant.
6. An information security management and monitoring system based on big data according to claim 5, characterized in that, The second evaluation unit is used to preset a first security level threshold X1 and a second security level threshold X2, and the first security level threshold X1 > the second security level threshold X2. Compare and evaluate the text sensitivity index MGbh with the first security level threshold X1 and the second security level threshold X2 respectively to determine the security level corresponding to the text sensitivity index MGbh of this first text, including: If the text sensitivity index MGbh < the second security level threshold X2, it indicates that the sensitivity of the first text is low, and the first security level is generated; If the second security level threshold X2 ≤ the text sensitivity index MGbh ≤ the first security level threshold X1, it indicates that the sensitivity of the first text is medium, and the second security level is generated; If the text sensitivity index MGbhi > the first security level threshold X1; it indicates that the sensitivity of the first text is high, and the third security level is generated.
7. An information security management and monitoring system based on big data according to claim 1, characterized in that, The data desensitization module includes a feature encoding unit and a desensitization unit; The feature encoding unit is used to generate corresponding desensitization strategies according to the first security level, the second security level, and the third security level and execute them, including: Generating a first desensitization strategy according to the first security level, including: converting 50% of the sensitive fields in the first text into an anonymous encoding format, specifically replacing 50% of the sensitive fields in the first text with the symbol "#” and retaining 50% of the original data; Generating a second desensitization strategy according to the second security level, including: converting 30% of the sensitive fields in the first text into an anonymous encoding format, specifically randomly generated fictional data, and covering 70% of the sensitive fields with black blocks; Generating a third desensitization strategy according to the third security level, including: performing the strictest desensitization process, covering 90% of the sensitive fields in the first text, converting the remaining 10% of the sensitive fields into an anonymous encoding format, specifically replacing them with the character "@”, and performing a hash process on the first text to generate a first key, and the first text can be read only after obtaining the resolution of the first key; The desensitization unit is used to count the first text converted by the feature encoding unit, convert the sensitive fields into non-sensitive fields to form a second text, and establish a desensitized data set.
8. An information security management and monitoring system based on big data according to claim 1, characterized in that, The sub-library management module includes a data partitioning unit, a distributed storage unit, an encryption processing unit, and a data analysis unit; The data partitioning unit is used to partition the desensitized data set into a first word library, a second word library, and a third word library according to the first security level, the second security level, and the third security level; The distributed storage unit distributes and stores the first word library, the second word library, and the third word library on different physical nodes; The encryption processing unit uses a combined algorithm of AES and RSA to encrypt the desensitized data sets of the first word library, the second word library, and the third word library, generates first encrypted data and stores it in the supervision database; The data analysis unit is used to extract the valid value features and redundant value features of the desensitized and encrypted desensitized data set, collect and obtain the user access frequency Fwpl for statistical analysis, and calculate and obtain the sensitive data protection coefficient Kp through the following formula: Where, V t represents the valid value of the desensitized data, V r represents the redundant value of the desensitized data, C represents the total number of desensitization information included, α represents the sensitive information weight coefficient, R represents the number of sensitive information that can be retrieved when restoring the data, and β represents the information recovery ability influence coefficient.
9. An information security management and monitoring system based on big data according to claim 1, characterized in that, The security prevention and control module is used to preset a security threshold Pc, and compare and evaluate the sensitive data protection coefficient Kp with the security threshold Pc to generate a security evaluation result, including: If the sensitive data protection coefficient Kp > the security threshold Pc, it indicates that the protection ability for the desensitized data set is qualified; If the sensitive data protection coefficient Kp ≤ the security threshold Pc, it means that the protection ability for the desensitized data set is unqualified, and a protection policy is generated, including: deleting the redundant value V of the desensitized data r , and performing secondary processing on the first encrypted data. The secondary processing includes splitting processing and additional encryption processing. The splitting processing includes splitting the first encrypted data into multiple data blocks, and applying an independent second key to each data block respectively. The second key is dynamically rotated after the user access frequency reaches 10 - 15 times; The additional encryption process includes: applying an additional encryption algorithm, Blowfish or ChaCha20, to the first encrypted data, and adding a hash checksum and a digital signature during the encryption process to generate the second encrypted data.
10. A big data-based information security management and monitoring system according to claim 9, characterized in that, The first encrypted data includes: user identity information, access log information, and data encryption meta-information; The data encryption meta-information includes the encryption algorithm type, key version, and encryption time; The second encrypted data includes: the result information after applying Blowfish or ChaCha20 to the first encrypted data, the independent key-related information generated for each data block, and the hash value and digital signature.
Citation Information
Cited By
Multi-region data compliance adaptation engine intelligent system
CN121167752A
A multi-regional data compliance adaptation engine intelligent system
CN121167752B