Data security usage control method based on customs data platform
By adopting the Presto distributed query engine and multi-source heterogeneous data source access technology on the customs data platform, combined with data hierarchical classification, encryption and audit functions, the problems of data silos and security management in the customs big data platform are solved, and the security control and traceability of the entire process of data use are achieved.
Patent Information
- Application Number
- CN202210823856.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The customs big data platform lacks data security use control methods that run through the entire process of data collection, data management, data processing and data services, making it difficult to communicate with data, forming an island, and lacks audit and traceability of the entire process of data use.
The data security use control method based on the customs data platform is adopted, and the access and metadata synchronization of multi-source heterogeneous data sources is performed by selecting and calling the Presto distributed query engine, data hierarchical classification and sensitive data encryption are realized, and combined with RBAC permission management and audit functions, authorization and auditing of the entire data use process is realized.
It realizes security control of the entire process of data use, ensures that the data is authorized and audited in the collection, processing and service process, achieves data traceability effect, and solves the serious problems of data silos and security management.
Smart Images

Figure CN115114645B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of customs data management, and in particular to a method for controlling the secure use of data based on a customs data platform. Background Art
[0002] Customs business data is the core data asset of the customs, which plays an important role in maintaining national economic and information security and is an important data resource of the country. With the accelerated development of new generation information technologies such as cloud computing and big data, the application scenarios of customs business data are more abundant, the application fields are wider, and at the same time, the data security management situation is more severe.
[0003] Due to the complexity of customs business and the large number of systems, each business department is relatively independent, the data is stored separately, and the perspectives on data are completely different. When the relevant systems were designed, no unified data storage method and format were determined, which ultimately led to difficulties in data interoperability and the formation of data islands.
[0004] Currently, the customs big data platform lacks an effective control method that runs through the entire process of data collection, data management, data processing, and data services, can authorize data with a minimized scope, audit the entire process of data use, and achieve the effect of data traceability. Therefore, this application proposes a new technical solution. Summary of the Invention
[0005] In order to closely match the actual situation of the application of customs business data and effectively control the security of the entire process of data use, this application provides a method for controlling the secure use of data based on a customs data platform.
[0006] This application provides a method for controlling the secure use of data based on a customs data platform, adopting the following technical solutions:
[0007] A method for controlling the secure use of data based on a customs data platform includes the following steps:
[0008] Q1. Select and call a pre-established Presto distributed query engine;
[0009] Q2. Based on the jdbc interface plug-in opened by the Presto distributed query engine itself as a development tool, perform secondary development to connect to multi-source heterogeneous data sources, including:
[0010] Select the target data source type; among them, the data source includes any one or more of Mysql, SqlServer, Oracle, Hive, GBase, and HBase;
[0011] Input the IP address and port of the database, the corresponding database name, and the instance name to complete the creation and configuration of the data mapping instance;
[0012] Add a Catalog module, which is used to provide a Rest service interface for loading configuration files; the server nodes of each Presto distributed query engine regularly request the Web service of the manager cluster, query the ES configuration library to obtain the configuration information of the existing target databases;
[0013] After the instance creation is completed, through the Catalog module, regularly obtain the existing information in the ES configuration library for comparison to verify the validity of the link; among them, the ES configuration library is a database storing database configurations and is used to ensure efficient reading of database configurations and synchronization of database information;
[0014] If the validity check of the link passes, the data instance information is sent through the Catalog module to synchronize the relevant data table metadata information under all or specified databases of the target data source to complete the metadata synchronization;
[0015] Q3. Data classification and grading;
[0016] Q4. Define sensitive data encryption policies;
[0017] Q5. Collect data, encrypt the data according to the encryption policy, and store it;
[0018] Q6. Data authorization, decrypt for use, and audit and store.
[0019] Optionally, the definition of sensitive data encryption policies includes: setting sensitive data encryption policies based on the classification and grading of data and the characteristics of field attribute values in data classification and grading; among them, the encryption policies include any one or more of sm2, aes, 3des, and sm4.
[0020] Optionally, encrypt the data during the data collection process, and the collected and stored data assets exist in three forms: data tables, data spaces, and data service interfaces.
[0021] Optionally, the data authorization, decrypt for use, and audit and store include: when the authorization object is a data table or a data space, when receiving user interaction information for data processing, read the level of the data table and the corresponding encryption policy, and decrypt and display the result data; and,
[0022] If the displayed content involves sensitive data fields, perform secondary processing and display using methods such as masking, rearrangement, or obfuscation, and record the user's data processing query statements and status.
[0023] Optionally, if the authorization object is a data service interface, when it is called by the front-end application, the DataService interface gateway reads the bound data space level and data encryption policy, and the SDK package referenced by the caller decrypts the transmitted data according to the corresponding policy.
[0024] Optionally, audit all Api access behaviors and conduct hierarchical audits according to the data classification of the data sources used by the interfaces;
[0025] If the data used for auditing includes two parts, one part is information such as access time, interface request, interface parameters, calling IP, authentication information, user name, role, organization, access success or failure, result code, return result details, number of returned data, size of returned data volume, returned data fields, traceable data query statements, etc. recorded in the interface service; the second part is the full amount of data generated by the interface; then execute the custom audit policy logic, which includes:
[0026] Set the fields to be audited according to the sensitivity level of the data;
[0027] Customize the audit data storage period according to the audit requirements of different data;
[0028] Among them, the audit of data is divided into three levels: level one only audits records, level two requires adding fields, and level three requires auditing all fields.
[0029] In summary, the present application includes at least one of the following beneficial technical effects: It can penetrate the entire process of data collection, data management, data processing, and data services, and can perform minimum-range authorization and full-process audit of data usage on data, achieving effective control of the data traceability effect. Brief Description of the Drawings
[0030] Figure 1 is a flowchart of the present application. Detailed Embodiment
[0031] The following Figure 1 further describes the present application in detail.
[0032] The embodiment of the present application discloses a method for controlling the secure use of data based on a customs data platform.
[0033] Referring to Figure 1 , the method for controlling the secure use of data based on a customs data platform includes at least two parts. One is the management of multi-source heterogeneous data assets; the other is to effectively and securely control the entire process of data usage in real time based on the first part.
[0034] For ease of understanding, first, an explanation of the management of multi-source heterogeneous data assets is given. Specifically:
[0035] Q1. Select and call the pre-established Presto distributed query engine;
[0036] Q2. Based on the jdbc interface plugin opened by the Presto distributed query engine itself as a development tool, perform secondary development.
[0037] The above is divided into:
[0038] Through secondary development based on the jdbc interface plugin opened by Presto itself, according to the communication messages of different data sources, connect to multi-source heterogeneous data sources such as Mysql, SqlServer, Oracle, Hive, GBase, and HBase; that is, transform the Presto source code, transform the mysql database module in the original Presto to provide more complete and diverse functions and expose them through Rest interfaces to improve availability and available scope.
[0039] Use a unified standard instance connection name to represent the unique data table path, providing a basis for realizing cross-database query and analysis of multi-source data.
[0040] By transforming the Presto component, add a Catalog module, which mainly provides a Rest service interface for loading configuration files.
[0041] Each Presto server node will regularly request the Web service of the manager cluster (the reason why the server node can request this is that the Presto server is an nginx server, and the management cluster uses nginx), query the ES configuration library to obtain the configuration information of the existing target database, so as to implement the authentication function for the database (table), and the permissions for querying, modifying, and deleting the data table are controlled separately.
[0042] In another embodiment of the present application, an audit function is also added to send audit data to the audit service ES cluster; that is, configure the audit logic to send audit data to the audit service ES cluster; among them, the audit data is generated from the data access authentication model; the ES cluster is a cluster formed by a series of related data storage tables, and at least one type of table is a table for storing audit data.
[0043] In one embodiment of the present application, specifically:
[0044] 1). Receive the interactive data input by the user, and select the target data source type on the corresponding page according to the instructions in the interactive data; among them, the data source includes any one or more of Mysql, SqlServer, Oracle, Hive, GBase, and HBase;
[0045] Identify interactive data, obtain the IP address and port of the database, the corresponding database name, and the instance name, and complete the creation and configuration of the data mapping instance.
[0046] 2) After completing the instance creation, through the Catalog module, regularly obtain the existing information in the ES configuration library for comparison to verify the validity of the link; among them, the ES configuration library is a database that stores database configurations and is used to ensure efficient reading of database configurations and synchronize database information.
[0047] If the validity check of the link passes, through the Catalog module, issue data instance information and synchronize the relevant data table metadata information under all or specified databases of the target data source.
[0048] It can be understood that this method is implemented in real time on the existing customs (big) data platform. Therefore, as described above, through the Catalog module, it can actually be understood that the platform or the system obtained by applying this method on the platform can execute subsequent actions through the Catalog module; the same applies to the following similar content and will not be elaborated.
[0049] 3) When the data instance information is issued through the Catalog module and the data instance connection is successfully created.
[0050] At this time, the administrator can view the data instance connection status in the "Physical Directory" of a pre-established data asset module on the data platform, manage the data assets and display the data structure, and monitor the status of the underlying physical database, tables, and fields in real time.
[0051] At the same time, the administrator can further set the method for dynamically confirming changes to the data source, specifically as follows:
[0052] For existing data instances, regularly obtain the relevant data table information, compare it with the existing information, and give prompts according to different comparison results; moreover, the data instances with changes are set to take effect only after receiving or obtaining the confirmation of the data administrator.
[0053] The above-mentioned prompts according to different comparison results include: if some data tables in the data source are deleted, the corresponding page shows that the table has been deleted (character information); if the table structure of the data table changes, the corresponding page shows that the table has been updated.
[0054] After completing the above-mentioned management of multi-source heterogeneous data assets, the control of the entire process of data collection, data management, data processing, and data services for the customs (big) data platform can be implemented. Specifically, this application also includes:
[0055] Q3. Data classification and grading;
[0056] Q4. Define the encryption policy for sensitive data;
[0057] Q5. Collect data, encrypt the data according to the encryption policy, and store it;
[0058] Q6. Authorize the data, decrypt and use it, and audit the storage.
[0059] The following explains Q3 - Q6 and other appendices:
[0060] For the data classification and grading in Q3, that is, after completing the metadata information synchronization operation of the target data source, the administrator can, on a pre - established "business directory" function page, uniformly manage the synchronized multi - source heterogeneous database and table information according to the business dimension and specific grading specifications.
[0061] That is, this application can be further set to: receive or obtain the administrator's instruction, establish a unified and maintainable data label directory, create data business labels under the corresponding directory according to the business characteristics, and add labels to the corresponding data tables to achieve the labeled management of data assets.
[0062] Define the encryption policy for sensitive data in Q4; for example: referring to the data security governance standard in the "Business Data Security Classification and Grading Standard", combined with the data classification and field attribute value characteristics in Q3, set the encryption policy for sensitive data, such as sm2, aes, 3des, and sm4, etc.
[0063] In an embodiment of this application, after completing the above - mentioned policy settings, the acquisition task of the target data can be configured. By using the selected data encryption policy, the data is encrypted and stored during the acquisition process; the acquired and stored data assets can exist in three forms: data tables, data spaces, and data service interfaces.
[0064] Subsequently, use the system unified authorization center based on the RBAC permission management model to authorize the synchronized and acquired multi - source heterogeneous data and data assets (data tables, data storage spaces, and data service interfaces) in combination with different user ranks and data usage requirements.
[0065] It should be noted that after this application is applied based on the existing customs data platform, the resulting system is called the customs data security control system, hereinafter referred to as the "system" for short.
[0066] In an embodiment of this application:
[0067] If the authorization object is a data table and a data space, when the user performs data processing (including operations such as statistical query analysis, etc.) according to specific permission restrictions, the system will read the level of the data table and the corresponding encryption policy, and decrypt and display the result data. If the displayed content involves sensitive data fields, secondary processing and display will be performed using methods such as masking, rearrangement, or obfuscation; at the same time, the system will record the user's data processing query statements and status.
[0068] If the authorization object is a data service interface, when the front-end application calls the interface, the DataService interface gateway of the system will automatically read the bound data space level and data encryption policy, and the SDK package referenced by the caller will decrypt the transmitted data according to the corresponding policy.
[0069] Furthermore, this application is set as follows: The system audits all Api access behaviors and conducts hierarchical audits according to the data classification of the data sources used by the interfaces. If the data used for auditing includes two parts, one part is the information such as access time, interface request, interface parameters, call IP, authentication information, user name, role, organization, access success or failure, result code, return result details, number of returned data, size of returned data volume, returned data fields, traceable data query statements, etc. recorded in the interface service. The second part is the full amount of data generated by the interface. In this case, the audit policy can be customized, including: 1) For some fields that need to be audited, they can be set specifically according to the sensitive level of the data; 2) According to the audit requirements of different data, the audit data storage period can be customized. Data audits are divided into three levels: level one only audits records, level two requires adding fields, and level three requires auditing all fields.
[0070] The following takes the application of the statistical analysis department of a certain customs that is currently in trial operation as an example for illustration:
[0071] Data administrator roles, data collection administrators, and application personnel are set for the system.
[0072] 1) The data administrator role is mainly responsible for data governance-related work (including: secondary processing of data assets, hierarchical classification management of data assets, and maintenance of business categories, and unified authorization of data assets on the big data platform (such as models, interfaces, data));
[0073] 2) The data collection administrator classifies and grades the metadata information of each source heterogeneous data after synchronization according to the "Customs Business Data Security Classification and Grading Standard". If there are sensitive fields such as ID numbers and enterprise numbers in the table, it can be determined as level 3 data, and a sensitive data encryption policy can be configured, such as selecting the sm4 algorithm for data encryption storage;
[0074] 3) After completing the classification and grading work, data collection tasks can be created for the target data source according to data collection requirements, data collection strategies can be configured, and daily data collection work can be completed regularly.
[0075] For example, for the list of cross-border import and export declaration enterprises, data collection must be executed once at 15 minutes, 30 minutes, 45 minutes, and on the hour every hour.
[0076] 4) After the data is collected, the data administrator authorizes the data assets according to specific individual or application needs.
[0077] For example, the application of a certain customs statistical analysis department will call the data interface published by the customs data security control system every 10 minutes to obtain the index result data calculated by the big data platform; for this scenario, the data administrator only needs to grant the application responsible person the access and call permissions for the relevant interface; usually, when the users of the statistical analysis department want to query and statistically analyze a certain level-3 data table according to work requirements, they only need to grant the "access" and "query" permissions of the table to the specified users.
[0078] 5) When the application personnel query and analyze the authorized data, the system will automatically record information such as the data script, user name, organization, role, data table accessed, access time, traffic, and number of records queried by the personnel. At the same time, the system automatically performs fuzzy processing on relevant sensitive fields. For example, for the ID number, the 4th to 8th digits in the middle are masked.
[0079] As can be seen from the above, this application is close to the actual situation of customs business data applications and can effectively and securely control the entire process of data use.
[0080] The above are all the preferred embodiments of this application. The protection scope of this application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of this application should be covered within the protection scope of this application.
Claims
1. A method for controlling the secure use of data based on a customs data platform, characterized in that, The following steps are involved: Q1. Select and call the pre-built Presto distributed query engine; Q2. Based on the open jdbc interface plug-in of the Presto distributed query engine itself as a development tool, secondary development is carried out to connect to multi-source heterogeneous data sources, including: Select the target data source type; data sources include any one or more of MySQL, SqlServer, Oracle, Hive, GBase, and HBase; Enter the database IP address and port, corresponding database name, and instance name to complete the creation and configuration of the data mapping instance; Add a Catalog module, which is used to provide a REST service interface for loading configuration files. The server nodes of each Presto distributed query engine periodically request the Web service of the manager cluster to query the ES configuration library to obtain the configuration information of the existing target database. After the instance is created, the Catalog module is used to periodically obtain the existing information in the ES configuration library for comparison and to verify the validity of the link. The ES configuration library is a database that stores database configurations and is used for efficient reading of database configurations to ensure database information synchronization. If the validity check of the link passes, the data instance information is sent through the Catalog module, and the metadata information of the relevant data tables in all or the specified database of the target data source is synchronized to complete the metadata synchronization; Q3. Data classification and grading; Q4. Define encryption strategy for sensitive data; Q5. Collect data, encrypt data according to encryption strategy, and store it; Q6. Data authorization, decryption, use, and audit storage; The definition of sensitive data encryption strategy includes: setting the encryption strategy of sensitive data based on the classification of data and field attribute value characteristics in data classification; wherein the encryption strategy includes any one or more of SM2, AES, 3DES, and SM4; The data is encrypted during the data collection process, and the collected and stored data assets exist in three forms: data table, data space, and data service interface; The data authorization, decryption and storage auditing include: when the authorization object is a data table or a data space, when receiving user interaction information for data processing, reading the level of the data table and the corresponding encryption strategy, and decrypting and displaying the result data; and, If the displayed content involves sensitive data fields, it will be processed and displayed again using masking, rearrangement or obfuscation methods, and the user's data processing query statement and status will be recorded.
2. The method for controlling the secure use of data based on a customs data platform according to claim 1, characterized in that: If the authorization object is a data service interface, when the front-end application calls it, the DataService interface gateway reads the bound data space level and data encryption policy, and the SDK package referenced by the caller decrypts the transmitted data according to the corresponding policy.
3. The method for controlling the secure use of data based on a customs data platform according to claim 1, characterized in that: Audit all API access behaviors and conduct hierarchical audits according to the data classification of the data source used by the interface; If the data used for auditing consists of two parts, one part is the access time, interface requests, interface parameters, calling IP, authentication information, username, role, organization, access success or failure, result code, return result details, number of returned data, size of returned data, returned data fields, and traceable data query statement information recorded in the interface service; The second part is the full amount of data generated by the interface; Then execute the custom audit policy logic, which includes: Set the required fields to be audited according to the sensitivity level of the data; Customize the audit data storage period according to the audit requirements of different data; Among them, the audit of data is divided into three levels: level one only audits records, level two requires adding fields, and level three requires auditing all fields.
Citation Information
Patent Citations
A heterogeneous data source visual query method
CN109815283A