Big data-based data security management and control method, device and equipment, and storage medium
By compressing, segmenting, feature-fusion, and classifying the target unit's data, and combining this with an identification model for security auditing, the problem of low data security has been solved, and efficient data security management has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2022-11-02
- Publication Date
- 2026-04-14
AI Technical Summary
In the era of big data and artificial intelligence, the business data of various departments are independent and the management standards are not uniform, resulting in low data security, difficulty in effectively identifying the authenticity of materials and protecting data, low targeting of existing systems, inability to directly verify data in other institutions' databases, and advanced material forgery techniques that are difficult to detect.
By acquiring the target unit's original data and pre-set authoritative third-party data, compressing and segmenting them, filtering and merging features and classifying them, using a pre-set identification model to conduct standardization checks and security audits, identifying the reliability of data sources, distinguishing between safe data and abnormal data, and managing them using standard or centralized control methods.
It improves data security and identification efficiency, ensures data standardization and reliability, and enables effective control over the original data of the target unit, thereby enhancing data security.
Smart Images

Figure CN115659401B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a data security management method, apparatus, electronic device, and computer-readable storage medium based on big data. Background Technology
[0002] In the era of big data and artificial intelligence, the importance of external data in the digital transformation of consumer finance services is becoming increasingly prominent. However, various problems arise when external data is actually implemented and applied, such as multiple departments importing data independently, inconsistent data management standards, and compromised data security.
[0003] Currently, the business data of various government agencies and public institutions are independent of each other. The agencies handling the business cannot directly view the original data in the databases of other agencies for verification. They can only rely on the seals and other evidence on the materials to judge their authenticity. Moreover, with the rapid development of network technology, the technology for forging materials is becoming increasingly sophisticated. Judging the authenticity of materials solely from the seals and other evidence is difficult for the staff of the receiving agency to verify, even if the materials have been forged or tampered with. Furthermore, since different departments need to protect different data, they generally use the same system to verify the authenticity and protect the data for convenience, which is not targeted enough. As a result, the identification results or security protection do not achieve the expected effect, leading to low data security. Summary of the Invention
[0004] This invention provides a data security management method, apparatus, and computer-readable storage medium based on big data, primarily aimed at addressing the problem of low data security. To achieve the above objective, this invention provides a data security management method based on big data, comprising:
[0005] Obtain raw data from the target unit and preset authoritative third-party data;
[0006] The original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain target data. Feature fusion is performed on the target data to obtain fused data, and the fused data is converted into a preset format and stored in a preset database.
[0007] Identify the business category of the preset third-party authoritative data, and classify the preset third-party authoritative data according to the business category to obtain classified data;
[0008] Using a preset first identification model, the fused data is subjected to standardization checks and security audits to obtain preliminary detection results;
[0009] The reliability of the classification data source is detected using a preset second identification model to obtain the third-party data reliability detection result;
[0010] When both the preliminary test results and the third-party data reliability test results are qualified, the security of the fused data is tested using a preset third identification model based on the classified data to obtain safe data and abnormal data. The safe data is managed in a standard management and control manner, and the abnormal data is managed in a centralized management and control manner.
[0011] Optionally, the step of compressing and segmenting the original data of the target unit, and filtering the segmented data to obtain the target data, includes:
[0012] The original data of the target unit is compressed to obtain compressed data;
[0013] The compressed data is segmented into multiple target data segments to be filtered.
[0014] Calculate the accuracy of each target data segment to be screened, and perform preliminary screening of the multiple target data segments based on the accuracy to obtain the initial screened data segments;
[0015] Calculate the verification value of the initial screening data segment based on the accuracy rate;
[0016] Based on the verification value and the preset verification threshold, determine whether the initial screening data segment is qualified;
[0017] When the initial screening data segment is unqualified, the target unit to which the unqualified initial screening data segment belongs is identified, the original data of the target unit is re-collected, and the step of compressing the original data of the target unit to obtain compressed data is returned until the initial screening data segment is qualified.
[0018] When the initial screening data segment is qualified, the qualified initial screening data segment is used as the target data.
[0019] Optionally, calculating the check value of the initial screening data segment based on the accuracy rate includes:
[0020] Extract the average feature values of the initial screening data segment and the target data segment to be screened;
[0021] The total amount of data in the initial screening data segment and the target data segment to be screened is calculated.
[0022] The verification value of the initial screening data segment is calculated based on the average feature value of the initial screening data segment, the average feature value of the target data segment to be screened, the total amount of data in the target data segment to be screened, the total amount of data in the initial screening data segment, and the accuracy of the initial screening data segment.
[0023] Optionally, the step of performing feature fusion on the target data to obtain fused data includes:
[0024] The target data is characterized by feature relationships using a pre-built feature relationship identification model to obtain the data features of the target data;
[0025] The data features are subjected to data structuring processing to obtain structured data groups;
[0026] Calculate the similarity between data in the structured data set;
[0027] Data with similarity values greater than a preset similarity threshold are fused to obtain fused data.
[0028] Optionally, the step of selecting a preset classification model to classify the preset third-party authoritative data according to the business category to obtain classified data includes:
[0029] Obtain multiple decision trees in the preset classification model, as well as the decision dimension index and decision conditions of at least one layer of nodes in each decision tree;
[0030] Based on the decision dimension index of the first node in the preset classification model, feature extraction is performed on the preset third-party authoritative data to obtain the feature value of the preset third-party authoritative data on the split dimension of the first node.
[0031] The feature value is judged based on the decision condition of the first node, and the second node to be traversed is determined from the branch nodes of the first node based on the judgment result.
[0032] Based on the current decision dimension index and decision conditions, continue to extract the feature values of the preset third-party authoritative data at the second node and determine the next node to be traversed until the decision tree traversal is completed, obtain the various categories of the preset third-party authoritative data, classify the preset third-party authoritative data according to each category, and obtain classified data.
[0033] Optionally, the step of using a preset second discrimination model to detect the reliability of the classification data source and obtain the third-party data reliability detection result includes:
[0034] Obtain the data source address corresponding to the classified data, use a preset second identification model to detect the accuracy of the data source address, and obtain the data source detection result;
[0035] When the data source detection result is qualified, the second identification model is used to detect the classification data for compliance detection, and a compliance detection result is obtained;
[0036] Based on the combined results of the data source testing and the compliance testing, the third-party data reliability testing results are obtained.
[0037] Optionally, the step of identifying the target unit to which the unqualified initial screening data segment belongs and re-collecting the original data of the target unit includes:
[0038] The security level of the target unit is identified according to preset rules;
[0039] Identify the data security level and data weight of the target unit to which the unqualified initial screening data segment belongs;
[0040] Query the total amount of all business data of the target unit and the data volume of the unqualified initial screening data segment in the corresponding target unit;
[0041] Calculate the impact of the unqualified initial screening data segment on the target unit to which it belongs;
[0042] When the influence degree is not greater than the preset influence degree threshold, the unqualified initial screening data segment is discarded;
[0043] When the impact exceeds the preset impact threshold, the business data of the target unit is re-collected and the original data of the target unit corresponding to the unqualified initial screening data segment is overwritten.
[0044] To address the aforementioned problems, the present invention also provides a data security management and control device based on big data, the device comprising:
[0045] The data fusion module is used to acquire the original data of the target unit and the preset third-party authoritative data; to compress and segment the original data of the target unit, and to filter the segmented data to obtain the target data; to perform feature fusion on the target data to obtain fused data; and to convert the fused data into a preset format and store it in a preset database.
[0046] The data classification module is used to identify the business category of the preset third-party authoritative data, select a preset classification model according to the business category to classify the preset third-party authoritative data, and obtain classified data.
[0047] The first identification module is used to perform standardization checks and security audits on the fused data using a preset first identification model to obtain preliminary detection results.
[0048] The second identification module is used to detect the reliability of the source of the classified data using a preset second identification model, and obtain the third-party data reliability detection result.
[0049] The third identification module is used to detect the security of the fused data based on the classification data and a preset third identification model when both the preliminary detection result and the third-party data reliability detection result are qualified, thereby obtaining safe data and abnormal data. The safe data is managed in a standard management manner, and the abnormal data is managed in a centralized management manner.
[0050] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:
[0051] At least one processor; and,
[0052] A memory communicatively connected to the at least one processor; wherein,
[0053] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the data security management method based on big data described above.
[0054] To address the aforementioned issues, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the aforementioned data security management method based on big data.
[0055] This invention compresses and segments the original data of the target unit, and filters the segmented data to obtain target data. Unqualified data in the original data is removed, resulting in higher data security. Feature fusion is performed on the target data to obtain fused data, reducing the data volume and improving the efficiency of subsequent data security identification. Furthermore, a classification model set is pre-constructed according to different business categories. A corresponding preset classification model can be selected based on the business category, facilitating accurate data classification according to different business categories. The preset authoritative third-party data is then classified to obtain categorized data, which is beneficial for subsequent data security detection of the fused data based on different categories of data. This results in higher accuracy in security detection and ensures data security. The system then performs a standardization check and security audit on the fused data using a first preset identification model to obtain preliminary detection results, ensuring the standardization and security of the fused data. A second preset identification model is used to detect the reliability of the source of the categorized data, obtaining third-party data reliability detection results to ensure the correctness and reliability of the third-party detection data. When both the preliminary detection results and the third-party data reliability detection results are qualified, a third preset identification model is used to detect the security of the fused data based on the categorized data, accurately distinguishing between secure data and abnormal data. Secure data is managed according to standard control methods, and abnormal data is managed according to centralized control methods, achieving effective control over the original data of the target unit and improving data security. Therefore, the data security management method, device, electronic device, and computer-readable storage medium based on big data proposed in this invention can solve the problem of low data security. Attached Figure Description
[0056] Figure 1 This is a flowchart illustrating a data security management method based on big data according to an embodiment of the present invention.
[0057] Figure 2 for Figure 1 The diagram shows a detailed implementation process for one step in a big data-based data security management method.
[0058] Figure 3 for Figure 1 The diagram shows a detailed implementation process for another step in the big data-based data security management method.
[0059] Figure 4 This is a functional block diagram of a data security management and control device based on big data provided in an embodiment of the present invention;
[0060] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the big data-based data security management method according to an embodiment of the present invention.
[0061] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0062] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0063] This application provides a data security management method based on big data. The executing entity of the data security management method based on big data includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the data security management method based on big data can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0064] Reference Figure 1 The diagram shown is a flowchart illustrating a data security management method based on big data according to an embodiment of the present invention. In this embodiment, the data security management method based on big data includes:
[0065] S1. Obtain the target unit's original data and preset authoritative third-party data.
[0066] In this embodiment of the invention, the target unit's original data refers to the original data that needs to be tested for data security, and the preset third-party authoritative data refers to authoritative data provided by national agencies and other platforms, such as data from credit reporting systems.
[0067] S2. Compress and segment the original data of the target unit, filter the segmented data to obtain target data, perform feature fusion on the target data to obtain fused data, and convert the fused data into a preset format and store it in a preset database.
[0068] For details, please refer to Figure 2 As shown, in S2, the original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain the target data, including:
[0069] S21. Compress the original data of the target unit to obtain compressed data;
[0070] S22. Divide the compressed data into multiple target data segments to be screened;
[0071] S23. Calculate the accuracy of each target data segment to be screened, and perform preliminary screening of the multiple target data segments to be screened based on the accuracy to obtain the initial screening data segments.
[0072] S24. Calculate the verification value of the initial screening data segment based on the accuracy rate;
[0073] S25. Based on the verification value and the preset verification threshold, determine whether the initial screening data segment is qualified;
[0074] When the initial screening data segment is unqualified, S26, identify the target unit to which the unqualified initial screening data segment belongs, re-collect the original data of the target unit, and return to step S21 until the initial screening data segment is qualified.
[0075] When the initial screening data segment is qualified, S27, the qualified initial screening data segment is used as the target data.
[0076] In this embodiment of the invention, the compressed data is divided into M target data segments to be screened, where M is a positive integer, and the accuracy of each target data segment to be screened is calculated using the following formula:
[0077]
[0078] Where η represents the accuracy of the target data segment to be screened, with a value range of (0,1); δ represents the precision coefficient of the target data segment to be screened; i represents the number of segments of the target data segment to be screened, with a value range of [1,M]; M represents the total number of segments of the target data segment to be screened; α i τ represents the total amount of data required for the i-th target data segment to be filtered; i γ represents the total amount of data in the i-th target data segment to be filtered; T represents the time taken to filter one of the target data segments to be filtered; f represents the filtering frequency of the target data segment to be filtered; γ represents the false positive rate of the target data segment to be filtered.
[0079] In this embodiment of the invention, when the verification value is greater than or equal to the preset verification threshold, the initial screening data segment is determined to be qualified; when the verification value is less than the preset verification threshold, the initial screening data segment is determined to be unqualified. The target unit to which the data in the unqualified initial screening data segment belongs is identified, the original data of the target unit is re-collected, and the step of compressing the original data of the target unit to obtain compressed data is returned until the initial screening data segment is qualified.
[0080] Further, S24 includes:
[0081] Extract the average feature values of the initial screening data segment and the target data segment to be screened;
[0082] The total amount of data in the initial screening data segment and the target data segment to be screened is calculated.
[0083] The verification value of the initial screening data segment is calculated based on the average feature value of the initial screening data segment, the average feature value of the target data segment to be screened, the total amount of data in the target data segment to be screened, the total amount of data in the initial screening data segment, and the accuracy of the initial screening data segment.
[0084] The check value of the initial screening data segment is calculated using the following formula:
[0085]
[0086] Where ψ represents the current verification value, with a value range of [0,1]; ζ represents the verification coefficient of the initial screening data segment; and represents the total amount of data in the target data segment to be screened. η represents the total amount of data in the initial screening data segment; η represents the accuracy of the initial screening data segment; λ represents the average feature value of the target data segment to be screened; and κ represents the average feature value of the initial screening data segment.
[0087] Furthermore, the process described in S26, which involves identifying the target unit to which the unqualified initial screening data segment belongs and re-collecting the original data of the target unit, includes:
[0088] The security level of the target unit is identified according to preset rules;
[0089] Identify the data security level and data weight of the target unit to which the unqualified initial screening data segment belongs;
[0090] Query the total amount of all business data of the target unit and the data volume of the unqualified initial screening data segment in the corresponding target unit;
[0091] Calculate the impact of the unqualified initial screening data segment on the target unit to which it belongs;
[0092] When the influence degree is not greater than the preset influence degree threshold, the unqualified initial screening data segment is discarded;
[0093] When the impact exceeds the preset impact threshold, the business data of the target unit is re-collected and the original data of the target unit corresponding to the unqualified initial screening data segment is overwritten.
[0094] In this embodiment of the invention, the influence of the unqualified initial screening data segment on the target unit is calculated using the following formula:
[0095]
[0096] Wherein, Ψ represents the data weight; S represents the unit security level; S ′ The data security level is indicated by Δa; the data capacity of the unqualified initial screening data segment in the corresponding target unit is indicated by Δa; and the total data volume is indicated by a.
[0097] In this embodiment of the invention, the data weight ranges from [0.1, 1].
[0098] In detail, the feature fusion of the target data described in S2 to obtain fused data includes:
[0099] The target data is characterized by feature relationships using a pre-built feature relationship identification model to obtain the data features of the target data;
[0100] The data features are subjected to data structuring processing to obtain structured data groups;
[0101] Calculate the similarity between data in the structured data set;
[0102] Data with similarity values greater than a preset similarity threshold are fused to obtain fused data.
[0103] In this embodiment of the invention, the pre-built feature relationship identification model is used to identify the feature relationships of the target data, and each feature relationship contains one or more data features. The pre-built feature relationship identification model can be a classification model built based on the BERT (Bidirectional Encoder Representation from Transformers) model.
[0104] In this embodiment of the invention, the preset database is a structured database, such as SQL.
[0105] In this embodiment of the invention, the fused data is data encrypted using the SM2 algorithm. During the encryption process, a digital certificate corresponding to the fused data is generated. Before performing standardization checks and security audits on the fused data using a preset first authentication model, the fused data is decrypted based on the digital certificate, thereby ensuring the data security of the fused data and preventing it from being tampered with during transmission.
[0106] In this embodiment of the invention, the original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain the target data. Unqualified data in the original data is deleted, which makes the data security higher. Feature fusion is performed on the target data to obtain fused data, which reduces the amount of data and makes the efficiency of subsequent data security identification higher.
[0107] S3. Identify the business category of the preset third-party authoritative data, and classify the preset third-party authoritative data according to the business category to obtain classified data.
[0108] In this embodiment of the invention, the data categories of the preset third-party authoritative data are first identified to obtain business categories. These business categories are data categories, such as insurance, banking, trust, and healthcare. Then, different classification models are selected based on different business categories to further classify the preset third-party authoritative data, resulting in categorized data. For example, the preset third-party authoritative data for the insurance category is further divided according to categories such as customer, product, agreement, and contract.
[0109] In this embodiment of the invention, the preset classification model can be a classification model constructed using a Random Forest (RF) model. The Random Forest model is a model that integrates multiple trees using the idea of ensemble learning, and its basic unit is a decision tree. Taking a classification problem as an example, each decision tree is a classifier. For an input sample, N trees will have N classification results. The Random Forest integrates all the classification voting results and designates the category with the most votes as the final output, thereby obtaining the optimal category.
[0110] For details, please refer to Figure 3 As shown in S3, the step of selecting a preset classification model based on the business category to classify the preset third-party authoritative data to obtain classified data includes:
[0111] S31. Obtain multiple decision trees in the preset classification model and the decision dimension index and decision conditions of at least one layer of nodes in each decision tree;
[0112] S32. Based on the decision dimension index of the first node in the preset classification model, perform feature extraction on the preset third-party authoritative data to obtain the feature value of the preset third-party authoritative data on the split dimension of the first node.
[0113] S33. Based on the decision conditions of the first node, the feature value is judged, and based on the judgment result, the second node to be traversed is determined from the branch nodes of the first node.
[0114] S34. Based on the current decision dimension index and decision conditions, continue to extract the feature values of the preset third-party authoritative data at the second node and determine the next node to be traversed until the decision tree traversal is completed, obtain the various categories of the preset third-party authoritative data, classify the preset third-party authoritative data according to each category, and obtain classified data.
[0115] In this embodiment of the invention, a storage layer is constructed according to the number of business categories, and the corresponding storage layer is randomly divided into regions to store the category data according to the type of the category data.
[0116] In this embodiment of the invention, a set of pre-built classification models is constructed according to different business categories. A corresponding preset classification model can be selected based on the business category, which facilitates accurate data classification according to different business categories. Classifying the preset authoritative third-party data to obtain categorized data is beneficial for subsequent security testing of the fused data based on different categories of data, thereby increasing the accuracy of security testing and ensuring data security.
[0117] S4. Using a preset first identification model, perform standardization checks and security audits on the fused data to obtain preliminary detection results.
[0118] In this embodiment of the invention, the data standardization detection is to detect whether the data type is correct, whether the data is reasonable, and whether it exceeds the lower limit of the array, etc., and the security audit is to perform security compliance detection on the fused data in accordance with relevant laws and regulations.
[0119] In this embodiment of the invention, the preset first discrimination model, the preset second discrimination model, and the preset third discrimination model are text matching models constructed from models such as BERT (Bidirectional Encoder Representation from Transformers) and RNN (Recurrent Neural Network).
[0120] In one embodiment of the present invention, the preset first identification model can be a text matching model constructed by the BERT model, which matches the fused data with the format of a preset data type to obtain a normative matching result; matches the fused data with relevant laws and regulations to obtain a security audit matching result; and combines the normative matching result and the security audit matching result to obtain a preliminary detection result.
[0121] S5. Use a preset second identification model to detect the reliability of the classification data source and obtain the third-party data reliability detection result.
[0122] In this embodiment of the invention, the compliance inspection is to protect data and ensure that sensitive data is not lost or damaged in accordance with the information inspection standards of international, national, governmental and organizational authorities, such as GDPR (General Data Protection Regulation) and PCI-DSS (Payment Card Industry Data Security Standard).
[0123] In one embodiment of the present invention, the preset second discrimination model can be a text matching model constructed by an RNN recurrent neural network. BLSTM (Bidirectional Long Short-Term Memory) is used to extract the data source address, information detection specification, and text features of the classified data corresponding to the classified data. Fully connected layers are used to calculate the matching scores between the text features of the classified data and the text features of the data source address and the text features of the information detection specification, respectively. The data source detection result and compliance detection result are obtained according to the matching score.
[0124] Specifically, S5 includes:
[0125] Obtain the data source address corresponding to the classified data, use a preset second identification model to detect the accuracy of the data source address, and obtain the data source detection result;
[0126] When the data source detection result is qualified, the second identification model is used to detect the classification data for compliance detection, and a compliance detection result is obtained;
[0127] Based on the combined results of the data source testing and the compliance testing, the third-party data reliability testing results are obtained.
[0128] In this embodiment of the invention, the classification data is data encrypted using the SM2 algorithm. During the encryption process, a digital certificate corresponding to the classification data is generated. Before using a preset second authentication model to detect the reliability of the source of the classification data, the classification data is decrypted based on the digital certificate, thereby ensuring the data security of the classification data.
[0129] In this embodiment of the invention, a preset second identification model is used to detect the accuracy of the data source address, ensuring the accuracy of the preset third-party authoritative data source and avoiding inaccurate subsequent security detection results due to an incorrect source of the preset third-party authoritative data.
[0130] S6. When both the preliminary detection result and the third-party data reliability detection result are qualified, the security of the fused data is detected using a preset third identification model based on the classified data to obtain safe data and abnormal data. The safe data is managed in accordance with the standard management method, and the abnormal data is managed in accordance with the centralized management method.
[0131] In this embodiment of the invention, the preset third discrimination model can be a text matching model constructed by models such as BERT (Bidirectional Encoder Representation from Transformers) or RNN (Recurrent Neural Network).
[0132] In this embodiment of the invention, the standard control method involves classifying the security data of the same batch and transmitting the corresponding security data to different preset first control nodes for storage based on the classification results; the centralized control method involves classifying the abnormal data and transmitting the classified data to different second control nodes for storage.
[0133] In this embodiment of the invention, the preset third identification model is used to further detect the security and compliance of the fused data based on the classification data to obtain secure data and abnormal data. The secure data is managed in a standard management manner, and the abnormal data is managed in a centralized management manner.
[0134] This invention compresses and segments the original data of the target unit, and filters the segmented data to obtain target data. Unqualified data in the original data is removed, resulting in higher data security. Feature fusion is performed on the target data to obtain fused data, reducing the data volume and improving the efficiency of subsequent data security identification. Furthermore, a classification model set is pre-constructed according to different business categories. A corresponding preset classification model can be selected based on the business category, facilitating accurate data classification according to different business categories. The preset authoritative third-party data is then classified to obtain categorized data, which is beneficial for subsequent data security detection of the fused data based on different categories of data. This results in higher accuracy in security detection and ensures data security. The method involves several steps. First, a pre-set first identification model is used to perform a standardization check and security audit on the fused data, obtaining preliminary detection results and ensuring the standardization and security of the fused data. A second pre-set identification model is then used to detect the reliability of the source of the categorized data, obtaining third-party data reliability detection results and ensuring the correctness and reliability of the third-party detection data. When both the preliminary detection results and the third-party data reliability detection results are qualified, a third pre-set identification model is used to detect the security of the fused data based on the categorized data, accurately distinguishing between safe and abnormal data. Safe data is managed according to standard control methods, and abnormal data is managed according to centralized control methods, achieving effective control over the original data of the target unit and improving data security. Therefore, the data security management method based on big data proposed in this invention can solve the problem of low data security.
[0135] like Figure 4 The diagram shown is a functional block diagram of a data security management and control device based on big data provided in an embodiment of the present invention.
[0136] The big data-based data security management device 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the big data-based data security management device 100 may include a data fusion module 101, a data classification module 102, a first authentication module 103, a second authentication module 104, and a third authentication module 105. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0137] In this embodiment, the functions of each module / unit are as follows:
[0138] The data fusion module 101 is used to acquire the original data of the target unit and the preset third-party authoritative data; to compress and segment the original data of the target unit, and to filter the segmented data to obtain the target data; to perform feature fusion on the target data to obtain fused data; and to convert the fused data into a preset format and store it in a preset database.
[0139] The data classification module 102 is used to identify the business category of the preset third-party authoritative data, select a preset classification model according to the business category to classify the preset third-party authoritative data, and obtain classified data.
[0140] The first identification module 103 is used to perform standardization checks and security audits on the fused data using a preset first identification model to obtain preliminary detection results;
[0141] The second identification module 104 is used to detect the reliability of the classification data source using a preset second identification model, and obtain the third-party data reliability detection result;
[0142] The third identification module 105 is used to detect the security of the fused data based on the classification data and a preset third identification model when both the preliminary detection result and the third-party data reliability detection result are qualified, thereby obtaining safe data and abnormal data, managing the safe data in a standard management method, and managing the abnormal data in a centralized management method.
[0143] In detail, the modules in the big data-based data security management and control device 100 described in this embodiment of the invention employ the same methods as described above. Figures 1 to 3 The data security management method based on big data described herein uses the same technical means and can produce the same technical effects, so it will not be elaborated here.
[0144] like Figure 5 The diagram shown is a structural schematic of an electronic device that implements a data security management method based on big data, according to an embodiment of the present invention.
[0145] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a data security management program based on big data.
[0146] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., executing data security management programs based on big data) and calls data stored in the memory 11 to perform various functions of the electronic device and process data.
[0147] The memory 11 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device, such as a plug-in portable hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory 11 can include both internal and external storage units of the electronic device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device, such as code for big data security management programs, but also to temporarily store data that has been output or will be output.
[0148] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 11 and at least one processor 10, etc.
[0149] The communication interface 13 is used for communication between the aforementioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, Bluetooth interface, etc.), typically used to establish communication connections between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), or, optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.
[0150] Figure 5 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 5 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0151] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0152] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0153] The data security management program based on big data stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When run in the processor 10, it can achieve the following:
[0154] Obtain raw data from the target unit and preset authoritative third-party data;
[0155] The original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain target data. Feature fusion is performed on the target data to obtain fused data, and the fused data is converted into a preset format and stored in a preset database.
[0156] Identify the business category of the preset third-party authoritative data, and classify the preset third-party authoritative data according to the business category to obtain classified data;
[0157] Using a preset first identification model, the fused data is subjected to standardization checks and security audits to obtain preliminary detection results;
[0158] The reliability of the classification data source is detected using a preset second identification model to obtain the third-party data reliability detection result;
[0159] When both the preliminary test results and the third-party data reliability test results are qualified, the security of the fused data is tested using a preset third identification model based on the classified data to obtain safe data and abnormal data. The safe data is managed in a standard management and control manner, and the abnormal data is managed in a centralized management and control manner.
[0160] Specifically, the specific implementation method of the processor 10 for the above instructions can be referred to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, and will not be repeated here.
[0161] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0162] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following:
[0163] Obtain raw data from the target unit and preset authoritative third-party data;
[0164] The original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain target data. Feature fusion is performed on the target data to obtain fused data, and the fused data is converted into a preset format and stored in a preset database.
[0165] Identify the business category of the preset third-party authoritative data, and classify the preset third-party authoritative data according to the business category to obtain classified data;
[0166] Using a preset first identification model, the fused data is subjected to standardization checks and security audits to obtain preliminary detection results;
[0167] The reliability of the classification data source is detected using a preset second identification model to obtain the third-party data reliability detection result;
[0168] When both the preliminary test results and the third-party data reliability test results are qualified, the security of the fused data is tested using a preset third identification model based on the classified data to obtain safe data and abnormal data. The safe data is managed in a standard management and control manner, and the abnormal data is managed in a centralized management and control manner.
[0169] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0170] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0171] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0172] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0173] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0174] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0175] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0176] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A data security management and control method based on big data, characterized in that, The method includes: Obtain raw data from the target unit and preset authoritative third-party data; The original data of the target unit is compressed and segmented, and the segmented data is filtered to obtain target data. Feature fusion is performed on the target data to obtain fused data, and the fused data is converted into a preset format and stored in a preset database. Identify the business category of the preset third-party authoritative data, select a preset classification model based on the business category, obtain various types of the preset third-party authoritative data by processing the preset classification model, and classify the preset third-party authoritative data according to each type to obtain classified data; The fused data is matched with the format of a preset data type using a preset first identification model to obtain a normative matching result. The fused data is then matched with relevant laws and regulations to obtain a security audit matching result. By combining the normative matching result and the security audit matching result, a preliminary detection result is obtained. The reliability of the classification data source is detected using a preset second identification model to obtain the third-party data reliability detection result; When both the preliminary test results and the third-party data reliability test results are qualified, the security of the fused data is tested using a preset third identification model based on the classified data to obtain safe data and abnormal data. The safe data is managed in a standard management and control manner, and the abnormal data is managed in a centralized management and control manner.
2. The data security management and control method based on big data as described in claim 1, characterized in that, The process of compressing and segmenting the original data of the target unit, and then filtering the segmented data to obtain the target data, includes: The original data of the target unit is compressed to obtain compressed data; The compressed data is segmented into multiple target data segments to be filtered. Calculate the accuracy of each target data segment to be screened, and perform preliminary screening of the multiple target data segments based on the accuracy to obtain the initial screened data segments; Calculate the verification value of the initial screening data segment based on the accuracy rate; Based on the verification value and the preset verification threshold, determine whether the initial screening data segment is qualified; When the initial screening data segment is unqualified, the target unit to which the unqualified initial screening data segment belongs is identified, the original data of the target unit is re-collected, and the step of compressing the original data of the target unit to obtain compressed data is returned until the initial screening data segment is qualified. When the initial screening data segment is qualified, the qualified initial screening data segment is used as the target data.
3. The data security management and control method based on big data as described in claim 2, characterized in that, The step of calculating the verification value of the initial screening data segment based on the accuracy rate includes: Extract the average feature values of the initial screening data segment and the target data segment to be screened; The total amount of data in the initial screening data segment and the target data segment to be screened is calculated. The verification value of the initial screening data segment is calculated based on the average feature value of the initial screening data segment, the average feature value of the target data segment to be screened, the total amount of data in the target data segment to be screened, the total amount of data in the initial screening data segment, and the accuracy of the initial screening data segment.
4. The data security management and control method based on big data as described in claim 1, characterized in that, The feature fusion of the target data to obtain fused data includes: The target data is characterized by feature relationships using a pre-built feature relationship identification model to obtain the data features of the target data; The data features are subjected to data structuring processing to obtain structured data groups; Calculate the similarity between data in the structured data set; Data with similarity greater than a preset similarity threshold are fused to obtain fused data.
5. The data security management and control method based on big data as described in claim 1, characterized in that, The processing based on the preset classification model yields various categories of preset third-party authoritative data, including: Obtain multiple decision trees in the preset classification model, as well as the decision dimension index and decision conditions of at least one layer of nodes in each decision tree; Based on the decision dimension index of the first node in the preset classification model, feature extraction is performed on the preset third-party authoritative data to obtain the feature value of the preset third-party authoritative data on the split dimension of the first node. The feature value is judged based on the decision condition of the first node, and the second node to be traversed is determined from the branch nodes of the first node based on the judgment result. Based on the current decision dimension index and decision conditions, continue to extract the feature values of the preset third-party authoritative data at the second node and determine the next node to be traversed until the decision tree traversal is completed, and obtain the various types of the preset third-party authoritative data.
6. The data security management and control method based on big data as described in claim 1, characterized in that, The step of using a preset second discrimination model to detect the reliability of the classification data source and obtaining the third-party data reliability detection result includes: Obtain the data source address corresponding to the classified data, use a preset second identification model to detect the accuracy of the data source address, and obtain the data source detection result; When the data source detection result is qualified, the second identification model is used to perform compliance detection on the classified data to obtain the compliance detection result; Based on the combined results of the data source testing and the compliance testing, the third-party data reliability testing results are obtained.
7. The data security management and control method based on big data as described in claim 2, characterized in that, The process of identifying the target unit to which the unqualified initial screening data segment belongs, and re-collecting the original data of the target unit, includes: The security level of the target unit is identified according to preset rules; Identify the data security level and data weight of the target unit to which the unqualified initial screening data segment belongs; Query the total amount of all business data of the target unit and the data volume of the unqualified initial screening data segment in the corresponding target unit; Calculate the impact of the unqualified initial screening data segment on the target unit to which it belongs; When the influence degree is not greater than the preset influence degree threshold, the unqualified initial screening data segment is discarded; When the impact exceeds the preset impact threshold, the business data of the target unit is re-collected and the original data of the target unit corresponding to the unqualified initial screening data segment is overwritten.
8. A data security management and control device based on big data, characterized in that, The device includes: The data fusion module is used to acquire the original data of the target unit and the preset third-party authoritative data; to compress and segment the original data of the target unit, and to filter the segmented data to obtain the target data; to perform feature fusion on the target data to obtain fused data; and to convert the fused data into a preset format and store it in a preset database. The data classification module is used to identify the business category of the preset third-party authoritative data, select a preset classification model according to the business category, obtain various types of the preset third-party authoritative data by processing the preset classification model, and classify the preset third-party authoritative data according to each type to obtain classified data. The first identification module is used to match the fused data with the format of a preset data type using a preset first identification model to obtain a normative matching result, match the fused data with relevant laws and regulations to obtain a security audit matching result, and combine the normative matching result and the security audit matching result to obtain a preliminary detection result. The second identification module is used to detect the reliability of the source of the classified data using a preset second identification model, and obtain the third-party data reliability detection result. The third identification module is used to detect the security of the fused data based on the classification data and a preset third identification model when both the preliminary detection result and the third-party data reliability detection result are qualified, thereby obtaining safe data and abnormal data. The safe data is managed in a standard management manner, and the abnormal data is managed in a centralized management manner.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the big data-based data security management method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data security management method based on big data as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data security centralized management and control method and system based on big data platform
CN112711757A