Sensitive data identification method and device based on big data and medium

Through the thermal update mechanism and metadata-driven policy selection, thread pool parallel processing and responsibility chain pattern, the problems of policy update and multi-type detection in sensitive data identification are solved, and flexible and efficient sensitive data identification and risk analysis are achieved.

CN120541101APending Publication Date: 2025-08-26INSPUR SMART TECH INNOVATION (SHANDONG) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510558073.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the identification of sensitive data, the existing technology has problems such as policy updates relying on downtime maintenance, difficulty in balancing efficiency and accuracy, weak multi-type detection synergy, and high system expansion and maintenance costs, making it difficult to achieve dynamic policy updates and optimize resource utilization.

Method used

Using a sensitive data identification method based on big data, the data sampling strategy and sensitive data identification method are dynamically loaded through the thermal update mechanism, the sampling strategy is selected based on metadata information, and the thread pool parallel processing and responsibility chain model is used to build a multi-type detection task chain, and a visual report is generated through sensitive rule base comparison.

Benefits of technology

It can respond to changes in business rules without shutdown during operation, improve flexibility and maintenance efficiency, shorten identification time, improve identification accuracy and credibility of results, and provide multi-dimensional risk analysis basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541101A_ABST
    Figure CN120541101A_ABST
Patent Text Reader

Abstract

The invention discloses a sensitive data identification method and device based on big data and a medium, and belongs to the technical field of big data processing. The method comprises the following steps: configuring a data sampling strategy and a sensitive data identification method; metadata information of the target data source is obtained, a data sampling strategy selection result is determined, and whether data initialization operation is started or not is judged; extracting data of the target data source according to a selected data sampling strategy to form an initial data set; generating a single-type sensitive data detection execution template and a multi-type sensitive data detection task chain, and performing sensitive data identification on the initial data set to generate an identification result; the recognition result is compared with a preset sensitive rule base to mark high-risk data, and a visual report is generated; and a data sampling strategy and a sensitive data identification method are adjusted through a hot update mechanism. By means of the method, dynamic strategy updating can be achieved to a certain extent, multi-type detection collaboration is improved, and the resource utilization rate is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of big data processing, and in particular to a sensitive data identification method, device and medium based on big data. Background Art

[0002] In the field of big data-driven sensitive data identification, existing technologies primarily perform detection based on static rule matching (e.g., regular expressions, keyword libraries) or single algorithm models (e.g., statistical learning, pattern recognition). To improve processing efficiency, some solutions introduce data sampling strategies (e.g., full scans, random sampling, or stratified sampling) to reduce the amount of data and shorten detection time. In the identification of multi-type sensitive data, common approaches use independent modules to process specific data types (e.g., ID numbers, phone numbers) in parallel and rely on strategy patterns to implement limited sampling algorithm switching. Existing technologies typically adjust detection logic through pre-configured rule bases and offline update mechanisms. Their core processes revolve around fixed rules and sampling parameters, such as defining regular expressions or adjusting sampling ratios through configuration files.

[0003] However, existing technologies have the following significant flaws: First, policy updates rely on downtime for maintenance and are unable to respond to dynamic business needs in real time. For example, adding new sensitive data types or adjusting sampling strategies requires interrupting system operation and redeploying code, resulting in impaired business continuity. Second, it is difficult to balance efficiency and accuracy. Although a full scan can cover all data, the processing time is too long; and random or stratified sampling may cause the omission of key sensitive information due to uneven data distribution. In particular, when there are large differences in field length and type, the sampling results are not representative enough and the missed detection rate increases significantly. Third, multi-type detection has weak coordination. When independent modules process different types of sensitive data, there is a lack of a unified scheduling mechanism, resulting in incomplete recognition of multiple types of sensitive information for the same data entry, and a lack of effective verification rules for conflicting results between modules (such as different modules labeling the same data with contradictory types). Fourth, the system expansion and maintenance costs are high, the code coupling is high, and the core logic needs to be modified to add new detection methods. The development cycle is long and prone to errors.

[0004] Therefore, how to achieve dynamic policy updates, improve the coordination of multi-type detection and optimize resource utilization has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present application provide a sensitive data identification method, device and medium based on big data to solve the following technical problems: how to achieve dynamic policy updates, improve the coordination of multi-type detection and optimize resource utilization.

[0006] In a first aspect, an embodiment of the present application provides a sensitive data identification method based on big data, which is applied to a big data processing system, wherein the big data processing system includes multiple policy processors, and is characterized in that the method includes: configuring a data sampling strategy and a sensitive data identification method, and importing them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategy includes at least a full extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification method includes a single type detection and a multi-type detection; obtaining metadata information of a target data source, and determining a data sampling strategy selection result based on the metadata information, and determining whether to start a data initialization operation; extracting data from the target data source according to the selected data sampling strategy to form an initial data set; wherein the data extraction strategy includes a full extraction strategy and a stratified sampling Extraction strategy, the initial data set includes the data entries to be detected and the associated metadata; build a policy factory, and match the data sampling strategy and the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template; wherein, the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; based on the preset chain of responsibility pattern, multiple policy processors are combined to build a multi-type sensitive data detection task chain; sensitive data is identified on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein, the identification result includes the sensitive data type and probability; the identification result is compared with the preset sensitive rule library to mark high-risk data and generate a visual report; the data sampling strategy and sensitive data identification method are adjusted through a hot update mechanism.

[0007] In one implementation of the present application, a data sampling strategy and a sensitive data identification method are configured and imported into a big data processing system through a preset hot update mechanism, specifically including: defining a strategy interface and declaring a data sampling strategy and a sensitive data identification method; implementing a specific strategy class for the strategy interface; wherein the specific strategy class includes at least a full scanning method, a layered sampling method, a regular expression matching algorithm, and a keyword matching method; encapsulating the specific strategy class into an independent file, and configuring an SPI service description file under the resource directory of the independent file to declare the strategy implementation class contained in the independent file; loading the independent file based on the preset ServiceLoader to parse the SPI service description file and instantiate the strategy implementation class; injecting the instantiated strategy implementation class into the memory of the big data processing system based on the preset Spring framework's automatic assembly mechanism to form an extensible strategy set; listening to external configuration change requests, loading the newly added or updated independent file through the preset independent class loader, and replacing the strategy implementation class with the same name in the strategy set to complete the hot update.

[0008] In one implementation of the present application, metadata information of a target data source is obtained, and whether to start a data initialization operation is determined based on the metadata information, specifically including: connecting to a target database of the target data source, and extracting metadata information of the target database, wherein the metadata information includes field name, field type, field length, and index attributes; judging whether the metadata information meets a preset sensitive data identification condition based on the field type and field length; wherein the sensitive data identification condition at least includes that the field length is greater than a preset minimum sensitive data length threshold; if the sensitive data identification condition is not met, the initialization operation of the metadata information is skipped; if the sensitive data identification condition is met, the data distribution characteristics of the metadata information are calculated, wherein the data distribution characteristics include data discreteness, data repetition rate, and time window distribution; determining a data sampling strategy selection result based on the data distribution characteristics, and generating an initialization operation start instruction; and storing the metadata information, the data sampling strategy selection result, and the initialization operation start instruction in a preset configuration center.

[0009] In one implementation of the present application, data is extracted from the target data source according to the selected data sampling strategy to form an initial data set, specifically including: if a full extraction strategy is selected, all data entries in the target data source are extracted to generate an initial data set; if a stratified sampling strategy is selected, the target data source is divided into multiple sub-data sources according to a time window or business rules, and the first number of data entries are extracted from each sub-data source according to a preset ratio; the extracted data entries are formatted and standardized to generate an initial data set.

[0010] In one implementation of the present application, a policy factory is constructed based on a policy pattern, and a target policy is matched from a data sampling policy and a sensitive data identification method to generate a single-type sensitive data detection execution template, specifically including: screening matching sensitive data identification methods from an extensible policy set according to field type; creating an execution template for the sensitive data identification method, defining data preprocessing, rule matching and result output processes; wherein data preprocessing includes format conversion and null value checking; associating the execution template with the thread pool, and adjusting the number of concurrent threads and task allocation strategy according to the initial data set size; obtaining key-value pair data, and distributing it to multiple threads according to a preset concurrency number to execute single-type sensitive data detection in parallel; marking the detection results as sensitive or non-sensitive, and recording the matching rule type and confidence level to the intermediate result cache to construct a single-type sensitive data detection execution template; wherein the intermediate result cache adopts a concurrent data structure.

[0011] In one implementation of the present application, multiple policy processors are combined based on a preset chain of responsibility pattern to construct a multi-type sensitive data detection task chain, specifically including: defining the interface of the chain of responsibility node, declaring the chain of responsibility node method and setting the next chain of responsibility node method; encapsulating the policy processor into a chain of responsibility node, and multiple chain of responsibility nodes; executing a specific type of sensitive data identification task in each chain of responsibility node; wherein the sensitive data identification task includes at least ID number detection, telephone number detection and bank card number detection; encapsulating the intermediate results between the chain of responsibility nodes as a context object and passing it to the next chain of responsibility node for incremental detection; recording the detection results of all chain of responsibility nodes based on a preset concurrent hash table, aggregating multiple types of results to generate a comprehensive sensitive data report, so as to construct a multi-type sensitive data detection task chain.

[0012] In one implementation of the present application, sensitive data identification is performed on an initial data set according to a single-type detection execution template and a multi-type detection task chain to generate an identification result, specifically including: calling the single-type detection execution template to perform parallel processing on the initial data set to generate a first sensitive data labeling result; inputting the first sensitive data labeling result into the multi-type detection task chain, and performing detection in the responsibility chain order of the multi-type detection task chain to update the sensitive data type and probability of the first sensitive data labeling result to generate a second sensitive data labeling result; performing a conflict check on the second sensitive data labeling result. If the same data entry is marked as sensitive by multiple responsibility chain nodes, the result with the highest confidence is taken as the identification result.

[0013] In one implementation of the present application, the identification results are compared with a preset sensitive rule library to mark high-risk data and generate a visual report, which specifically includes: determining the sensitive rule library; wherein the sensitive rule library includes a predefined sensitive data type and risk level mapping table; querying the risk level mapping table according to the sensitive type in the identification result to determine the risk level; if the risk level is higher than a preset threshold, marking the identification result as high-risk data; integrating the high-risk data and risk level into a structured report; and generating a visual report based on the structured report.

[0014] In a second aspect, an embodiment of the present application further provides a sensitive data identification device based on big data, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategies include at least a full-scale extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; obtain metadata information of the target data source, and determine the data sampling strategy selection result based on the metadata information, and determine whether to start the data initialization operation; extract data from the target data source according to the selected data sampling strategy to form an initial data The invention relates to a method for detecting sensitive data of a specific type and a sensitive data set; wherein, the data extraction strategy includes a full extraction strategy and a stratified sampling extraction strategy, and the initial data set includes the data entries to be detected and the associated metadata; a policy factory is constructed, and the target strategy in the data sampling strategy and the sensitive data identification method is matched to generate a single-type sensitive data detection execution template; wherein, the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; multiple policy processors are combined based on a preset chain of responsibility pattern to construct a multi-type sensitive data detection task chain; sensitive data is identified on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein, the identification result includes the sensitive data type and probability; the identification result is compared with the preset sensitive rule library to mark high-risk data and generate a visual report; the data sampling strategy and the sensitive data identification method are adjusted through a hot update mechanism.

[0015] On the third aspect, the embodiment of the present application also provides a non-volatile computer storage medium for sensitive data identification based on big data, which stores computer executable instructions, characterized in that the computer executable instructions are set to: configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategies include at least a full-scale extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; obtain metadata information of the target data source, and determine the data sampling strategy selection result based on the metadata information, and judge whether to start the data initialization operation; extract data from the target data source according to the selected data sampling strategy to form an initial data set; wherein the data extraction strategies include a full-scale extraction strategy and a stratified extraction strategy Sampling extraction strategy, the initial data set includes the data entries to be detected and the associated metadata; build a policy factory, and match the data sampling strategy and the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template; wherein, the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; based on the preset chain of responsibility pattern, multiple policy processors are combined to build a multi-type sensitive data detection task chain; sensitive data is identified on the initial data set according to the single-type detection execution template and the multi-type detection task chain, and an identification result is generated; wherein, the identification result includes the sensitive data type and probability; the identification result is compared with the preset sensitive rule library to mark high-risk data, and a visual report is generated; the data sampling strategy and sensitive data identification method are adjusted through a hot update mechanism.

[0016] The embodiments of the present application provide a method, device, and medium for identifying sensitive data based on big data. This application has at least the following technical effects: A pre-set hot update mechanism dynamically loads and replaces data sampling strategies and sensitive data identification methods, enabling real-time policy updates and expansions. This allows for responses to business rules or regulatory changes without downtime, significantly improving flexibility in responding to dynamic needs and maintenance efficiency.

[0017] Intelligently selects full or stratified sampling strategies based on metadata information to reduce redundant data processing. Combined with a thread pool to parallelize detection tasks, this significantly reduces the overall time required to identify sensitive data and improves throughput in large data volume scenarios.

[0018] By generating single-type detection templates through a policy factory and combining them with the chain of responsibility model to connect multiple detection nodes, the sensitive data identification process is modularized and organized. The context object transfer and concurrent hash table recording mechanism within the chain of responsibility effectively support the aggregation and conflict verification of multiple detection results, enhancing recognition accuracy and result credibility in complex scenarios.

[0019] The system compares the identification results with the sensitive rule base, automatically marks high-risk data, and presents the detection results intuitively through visual charts. This provides data governance personnel with a multi-dimensional risk analysis basis, ensuring the accurate formulation and rapid response of data security policies to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a sensitive data identification method based on big data provided in an embodiment of the present application; Figure 2 A schematic diagram of the internal structure of a sensitive data identification device based on big data provided in an embodiment of the present application.

[0021] For example To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] The embodiments of the present application provide a sensitive data identification method, device and medium based on big data to solve the following technical problems: how to achieve dynamic policy updates, improve the coordination of multi-type detection and optimize resource utilization.

[0023] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0024] Figure 1 This is a flowchart of sensitive data identification based on big data provided by the embodiment of this application. Figure 1 As shown, the embodiment of the present application provides a sensitive data identification method based on big data, which specifically includes the following steps: Step 1: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through the preset hot update mechanism; the data sampling strategies include at least full-scale extraction strategies and stratified sampling extraction strategies, and the sensitive data identification methods include single-type detection and multi-type detection.

[0025] By defining a policy interface, implementing a specific policy class, encapsulating it in a separate file, and leveraging the SPI mechanism and Spring framework's auto-assembly capabilities, data sampling policies and sensitive data identification methods can be dynamically loaded and updated. The hot update mechanism allows the embodiment of this application to dynamically replace policy implementation classes at runtime without requiring downtime.

[0026] Step 1.1. Define the policy interface and declare the data sampling strategy and sensitive data identification method.

[0027] Policy interface: A standardized interface used to constrain the specific implementation logic of data sampling strategies and sensitive data identification methods.

[0028] Data sampling strategy: rules for extracting part of the data from a large data set, such as full extraction or stratified sampling.

[0029] Sensitive data identification methods: Algorithms or rules used to detect sensitive information in data, such as regular expression matching or keyword matching.

[0030] For example, define two strategy interfaces: Sampling strategy interface: declares the executeSampling() method, which is used to extract data according to specific rules.

[0031] Identification method interface: declares the detectSensitiveData() method to execute sensitive data detection logic.

[0032] Step 1.2: Implement a specific strategy class for the strategy interface; wherein the specific strategy class includes at least a full scanning method, a stratified sampling method, a regular expression matching algorithm, and a keyword matching method.

[0033] Specific strategy class: A class that implements the strategy interface and contains specific algorithm logic. For example, the full scan class implements full data extraction logic.

[0034] For example, the full scan class overrides the executeSampling() method to extract all data from the data source.

[0035] Stratified sampling class: override the executeSampling() method to extract data in layers according to time or business rules.

[0036] Regular expression matching class: Override the detectSensitiveData() method to match sensitive data using a preset regular expression (such as ID card number or mobile phone number).

[0037] Keyword matching class: Override the detectSensitiveData() method to match sensitive data through the keyword library (such as "confidential" and "password").

[0038] Step 1.3: Encapsulate the specific policy class into an independent file, and configure the SPI service description file in the resource directory of the independent file to declare the policy implementation class contained in the independent file.

[0039] Independent file: a policy implementation class file encapsulated in a JAR package.

[0040] SPI service description file: A configuration file located in the resource directory of the JAR package, declaring the specific class name that implements the strategy interface in the JAR package.

[0041] For example, compile the full scan class, stratified sampling class, etc. into a JAR package.

[0042] Create a file in the META-INF / services / directory of the JAR package, name it the fully qualified name of the strategy interface (such as com.example.SamplingStrategy), and use the fully qualified name of the specific strategy class (such as com.example.FullSampling) as the file content.

[0043] Step 1.4: Load the independent file based on the preset ServiceLoader to parse the SPI service description file and instantiate the policy implementation class.

[0044] ServiceLoader: Java's built-in service loading tool, which dynamically loads implementation classes through the SPI mechanism.

[0045] For example, use ServiceLoader.load(SamplingStrategy.class) to load all classes that implement the sampling strategy interface.

[0046] Traverse the loaded classes and instantiate them to generate a set of strategy instances.

[0047] Step 1.5: Based on the preset Spring framework's automatic assembly mechanism, the instantiated strategy implementation class is injected into the memory of the big data processing system to form an extensible strategy set.

[0048] Spring auto-assembly: The Spring framework automatically injects Bean instances into the context of the big data processing system by scanning the class path.

[0049] For example, to enable component scanning in the Spring configuration file ( <context:component-scan>).

[0050] Add the @Component annotation to the strategy implementation class, and Spring will automatically inject it into the in-memory strategy collection.

[0051] Step 1.6: Listen for external configuration change requests, load the newly added or updated independent files through the preset independent class loader, and replace the policy implementation class with the same name in the policy set to complete the hot update.

[0052] Independent class loader: used to dynamically load new or updated JAR files to avoid conflicts with existing system classes.

[0053] For example, listen to change events of configuration centers (such as ZooKeeper).

[0054] When a new JAR file is detected, a separate class loader is used to load the file and parse the SPI description.

[0055] Replace the strategy implementation class with the same name in the strategy collection, and the old class instance is recycled by the garbage collection mechanism.

[0056] In a specific example, Company A's big data system needed to dynamically update its sensitive data identification methods to comply with newly enacted privacy regulations. The original system used regular expressions to match mobile phone numbers, but the new regulations required the addition of "virtual number" identification. Company A followed the following steps: Implement a new strategy class: Develop a "virtual number identification class", override the detectSensitiveData() method, and add a new regular expression for virtual numbers (such as 1\d{2}\*\*\*\d{4}).

[0057] Package as a JAR package: Package the new class as virtual-number-detector.jar and declare the class name in its SPI file.

[0058] Hot update: Upload the JAR package to the configuration center. After the system detects the change, it loads the package through an independent class loader, replacing the original identification method.

[0059] Verification effect: The system automatically uses the new policy to detect virtual numbers without restarting the service, and the old policy can still be rolled back.

[0060] Results: Through the hot update mechanism, Company A completed the policy upgrade within 10 minutes, avoiding business interruption and ensuring compliance with the latest regulatory requirements.

[0061] Step 2: Obtain metadata information of the target data source, determine the data sampling strategy selection result based on the metadata information, and determine whether to start the data initialization operation.

[0062] By analyzing the metadata of the target data source (such as field type, length, etc.), it is determined whether the sensitive data identification conditions are met, and the optimal sampling strategy is dynamically selected based on the data distribution characteristics, ultimately deciding whether to trigger the data initialization operation.

[0063] Step 2.1: Connect to the target database of the target data source and extract metadata information of the target database, where the metadata information includes field name, field type, field length, and index attributes.

[0064] Target database: The data storage system to be processed, such as a relational database (MySQL) or a NoSQL database (MongoDB).

[0065] Metadata information: describes the properties of the data structure, including field name, field type (such as string, number), field length (such as character limit), index properties (such as primary key, unique index), etc.

[0066] For example, a connection to a target database is established through the JDBC (Java Database Connectivity) protocol, and a standard SQL query (such as DESCRIBE table_name) is executed to obtain table structure information.

[0067] Parse the query results and extract metadata information for all fields in the target table, for example: Field name: user_address Field type: VARCHAR (variable-length character string) Field length: 255 Indexed properties: No index Step 2.2: Determine whether the metadata information meets the preset sensitive data identification conditions based on the field type and field length; wherein the sensitive data identification conditions at least include that the field length is greater than the preset minimum sensitive data length threshold.

[0068] Sensitive data identification conditions: Preset rules used to filter fields that may contain sensitive data. For example: The field type is string (VARCHAR, TEXT).

[0069] The field length must be greater than or equal to the minimum sensitive data length threshold (for example, a mobile phone number must be ≥11 digits).

[0070] For example, traverse the metadata information of all fields and filter out fields whose field type is string.

[0071] For the filtered fields, check whether their length is greater than or equal to the preset threshold (for example, if the field length is <11, the mobile phone number cannot be stored and is skipped directly).

[0072] Step 2.3: If the sensitive data identification conditions are not met, the initialization operation of the metadata information is skipped.

[0073] If a field does not meet the conditions (for example, the field type is a numeric value or the length is insufficient), the system records a log and marks the field as "no processing required", and it will no longer be involved in sensitive data identification in subsequent processes.

[0074] Step 2.4: If the sensitive data identification conditions are met, the data distribution characteristics of the metadata information are calculated, where the data distribution characteristics include data dispersion, data repetition rate, and time window distribution.

[0075] Data distribution characteristics: Statistical indicators that reflect the distribution patterns of data, including: Data dispersion: The proportion of different values ​​(for example, there are 80 different addresses in 100 data).

[0076] Data repetition rate: The ratio of repeated values ​​to the total data (for example, 80% of the data are repeated addresses).

[0077] Time window distribution: The distribution of data by timestamp (for example, data is concentrated in the past 3 months).

[0078] For example, data dispersion: execute the SQL query SELECTCOUNT(DISTINCTfield_name)FROMtable_name to calculate the proportion of unique values.

[0079] Data duplication rate: Execute SELECT(COUNT(*)-COUNT(DISTINCTfield_name)) / COUNT(*)FROMtable_name to calculate the duplication ratio.

[0080] Time window distribution: If the field contains a timestamp, execute SELECT MIN(time_field), MAX(time_field) FROM table_name to analyze the time span.

[0081] Step 2.5: Determine the data sampling strategy selection result based on the data distribution characteristics and generate the initialization operation start instruction.

[0082] For example: High repetition rate (>70%): Select a stratified sampling strategy to extract representative data by grouping the replicate values.

[0083] Concentrated time window (e.g., data from the past three months accounts for 90%): Select a time window sampling strategy to prioritize recent data.

[0084] Data dispersion is high (unique values ​​account for > 90%): Select a full scan strategy to ensure that all possible situations are covered.

[0085] Step 2.6: Store the metadata information, data sampling strategy selection results, and initialization operation start instructions in a preset configuration center.

[0086] Configuration Center: Centrally manages system configuration components (such as ZooKeeper and Nacos), used to store metadata, policy selection results, etc.

[0087] For example: Encapsulates metadata information (field name, type), strategy selection results (such as "stratified sampling"), and initialization instructions (such as "start") in JSON format.

[0088] Use the configuration center's API to write data to a predefined path (such as / config / data_init / user_table).

[0089] In a specific example, e-commerce platform B needs to identify sensitive data in the "address" field in the user table. The system performs the following process: Extract metadata: Connect to the database and obtain the "address" field type, which is VARCHAR, length 255, and no index.

[0090] Judgment condition: The field type is string and the length is ≥11 (meets the condition).

[0091] Calculate data distribution: Data dispersion: There are 95,000 different values ​​in 100,000 addresses (dispersion 95%).

[0092] Time window distribution: 80% of the data was added in the past month.

[0093] Select a sampling strategy: Because the time windows are concentrated, select "Time Window Stratified Sampling" to prioritize data from the past month.

[0094] Storage configuration: Writes the strategy selection results ("time window stratified sampling") and startup instructions to the configuration center.

[0095] Step 3: Extract data from the target data source according to the selected data sampling strategy to form an initial data set; wherein, the data extraction strategy includes a full extraction strategy and a stratified sampling extraction strategy.

[0096] According to the data sampling strategy determined in step 2 (full extraction or stratified sampling), extract data from the target data source and standardize the extracted data to generate a unified initial data set to provide input for subsequent sensitive data identification.

[0097] Step 3.1: If the full extraction strategy is selected, all data entries in the target data source are extracted to generate the initial data set.

[0098] Full extraction strategy: extracts all data from the target data source without any filtering or sampling.

[0099] Data entry: A single record in the target data source (such as a row of data in a database table).

[0100] For example, to perform a full query based on the target data source type (such as a database table or file): Database scenario: Execute the SQL statement SELECT * FROM table_name to obtain all data in the table.

[0101] File scenario: Read all the contents of a file (such as CSV, JSON).

[0102] Map the query results directly into a structured data set to generate the initial data set.

[0103] Step 3.2: If the stratified sampling strategy is selected, the target data source is divided into multiple sub-data sources according to the time window or business rules, and the first number of data entries are extracted from each sub-data source according to a preset ratio.

[0104] Stratified sampling strategy: Divide the data into multiple subsets (layers) according to specific dimensions (such as time windows, business rules), and then extract data proportionally from each layer.

[0105] Time window: divide the data by time range (such as by month or by day).

[0106] Business rules: Divide data by business attributes (such as user type and region).

[0107] For example, to divide the sub-data source: Time window division: Divide data into multiple time periods (for example, the past month, 1-3 months, and more than 3 months) based on a time field (such as create_time).

[0108] Business rule classification: Classification by field value (for example, user levels are "VIP" and "ordinary user").

[0109] Extract data proportionally: Set a sampling ratio for each sub-data source (for example, data from the past month accounts for 70%, and the first 1,000 records are sampled).

[0110] Execute hierarchical queries (such as SELECT * FROM table_name WHERE create_time >= '2023-01-01' LIMIT 1000).

[0111] Step 3.3: Standardize the format of the extracted data entries to generate the initial data set.

[0112] Format standardization: Unify data formats, eliminate redundant or conflicting data structures, and ensure data consistency.

[0113] For example, data cleaning: Remove empty value fields or fill in default values ​​(such as replacing NULL with "N / A").

[0114] Merge duplicate fields (e.g. merge "mobile number" and "contact number" into "phone_number").

[0115] Format conversion: Standardize date formats (e.g. convert "2023 / 01 / 01" to "2023-01-01").

[0116] Unify the encoding format (such as unifying the text encoding to UTF-8).

[0117] Data Validation: Check that the field type meets expectations (for example, whether a numeric field contains non-numeric characters).

[0118] In a specific example, Bank C needs to identify sensitive data in the "Transaction Notes" field in the transaction record table. The system operates according to the following process: Select a stratified sampling strategy: Divide the time window: Divide the data into three subsets by month: "January 2023", "February 2023", and "March 2023".

[0119] The sampling ratio is set as follows: data from the past month (March) accounts for 60%, and the first 5,000 records are selected; data from January and February each accounts for 20%, and the first 2,000 records are selected for each.

[0120] Perform stratified sampling: Execute SELECT * FROM transaction WHERE month = 3 LIMIT 5000 for the March data.

[0121] Run similar queries for the January and February data.

[0122] Format standardization: Clean data: Delete meaningless symbols (such as "***") in the "Transaction Notes" field.

[0123] Unified format: Convert the date field "transaction_date" to the "YYYY-MM-DD" format.

[0124] Generate the initial data set: Merge the cleaned 9,000 data items (5,000 + 2,000 + 2,000) into the initial data set.

[0125] Step 4: Build a policy factory and match the data sampling strategy with the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template. The single-type sensitive data detection execution template processes tasks in parallel through a thread pool.

[0126] Through the policy factory, data sampling strategies and sensitive data identification methods are dynamically bound to build a standardized detection process template, and the thread pool is used to achieve task parallelism to improve processing efficiency in big data scenarios.

[0127] Step 4.1: Filter matching sensitive data identification methods from the extensible policy set based on field type.

[0128] Policy Collection: Contains all loaded sensitive data identification methods (such as regular expression matching and keyword matching).

[0129] Field type matching: Select the appropriate recognition method based on the field type (e.g., string, numeric). For example, mobile phone number recognition only requires a string field.

[0130] For example: Traverse the field types of the initial dataset and filter the matching identification methods: String field: Use regular expression matching (such as ID number, mobile phone number) or keyword matching (such as "password" and "key").

[0131] Numeric fields: Call range matching (such as amount exceeding a threshold) or format matching (such as bank card number rules).

[0132] If there is no matching method for a field, it is marked as "No need to check".

[0133] Step 4.2: Create an execution template for the sensitive data identification method, defining the data preprocessing, rule matching, and result output processes; data preprocessing includes format conversion and null value verification.

[0134] Execution template: Standardized detection process, including three stages: data preprocessing, rule matching, and result output.

[0135] Data preprocessing: Ensure that the input data meets the detection requirements, such as unified format and null value processing.

[0136] For example: Data preprocessing: Format conversion: Convert date fields to ISO standard format (such as YYYY-MM-DD).

[0137] Empty value validation: filter empty value fields or fill them with default values ​​(such as "unknown").

[0138] Rule matching: Call the filtered recognition method (such as regular expression matching mobile phone numbers).

[0139] Result output: Generates structured results, including field name, detection result (sensitive / non-sensitive), matching rule type, and confidence level (such as 90% match).

[0140] Step 4.3: Associate the execution template with the thread pool and adjust the number of concurrent threads and task allocation strategy based on the initial dataset size.

[0141] Thread pool: manages a resource pool of concurrent threads to avoid the overhead of frequent creation / destruction of threads.

[0142] Number of concurrent threads: Dynamically adjusted based on the size of the dataset and system resources, for example, the number of CPU cores × 2.

[0143] For example: Initialize the thread pool: set the number of core threads (such as 10), the maximum number of threads (such as 50), and the task queue capacity (such as 1000).

[0144] Dynamic adjustment strategy: Small data sets (<10,000 records): Single-threaded sequential processing.

[0145] Large data sets (≥10,000 records): Allocate threads based on the amount of data (e.g., 1 thread for every 5,000 records).

[0146] Task allocation: Split the initial dataset into multiple subsets, and submit each subset as a task to the thread pool.

[0147] Step 4.4: Obtain key-value pair data and distribute it to multiple threads according to the preset concurrency number to perform single-type sensitive data detection in parallel.

[0148] Key-value pair data: a data unit organized in the form of key (unique identifier of data) and value (content of the field to be detected).

[0149] For example: Data chunking: Split the initial data set into multiple sub-chunks (e.g., 1,000 records / chunk) based on the preset concurrency number.

[0150] Task submission: Each sub-block is encapsulated as an independent task and submitted to the thread pool queue.

[0151] Parallel execution: The thread pool automatically allocates idle threads to execute tasks, and the detection logic calls the rule matching process in the execution template.

[0152] Step 4.5: Mark the detection result as sensitive or non-sensitive, and record the matching rule type and confidence level in the intermediate result cache to build a single-type sensitive data detection execution template; the intermediate result cache adopts a concurrent data structure.

[0153] Intermediate result cache: A thread-safe storage structure (such as ConcurrentHashMap) that temporarily stores detection results.

[0154] Confidence: An indicator of the strength of a rule match (such as the character coverage of a regular expression match).

[0155] For example: Result Tags: Sensitive data: Mark the field as "Sensitive" and record the matching rule type (such as "Mobile phone number regular expression matching").

[0156] Non-sensitive data: Mark as "non-sensitive" and record the reason (such as "no matching rule").

[0157] Cache persistence: Periodically write cache results in batches to a database or file to avoid memory overflow.

[0158] In a specific example, Medical System D needs to detect whether the "Diagnosis Description" field in a patient's medical record contains sensitive information (such as the name of a disease). The system operates according to the following process: Screening and identification methods: Set the field type to string and select keyword matching (the keyword library includes disease terms such as "cancer" and "HIV").

[0159] Create an execution template: Preprocessing: Remove special symbols (such as "*") in the description.

[0160] Rule matching: Compare the keyword library and mark it as sensitive if it matches.

[0161] Configure the thread pool: The initial data set contains 50,000 records, and the number of concurrent threads is set to 10 (each thread processes 5,000 records).

[0162] Parallel detection: The thread pool allocates tasks, and 10 threads execute keyword matching in parallel.

[0163] Record the results: 1,200 records containing the keyword "cancer" were detected with a confidence level of 100%. The results were cached and marked as sensitive.

[0164] Step 5: Combine multiple policy processors based on the preset responsibility chain model to build a multi-type sensitive data detection task chain.

[0165] Through the chain of responsibility pattern, multiple sensitive data identification tasks are connected in series into an ordered chain. Each node processes a specific type of detection task and passes intermediate results through context objects. Finally, multiple types of detection results are aggregated to generate a comprehensive report.

[0166] Step 5.1. Define the interface of the responsibility chain node, declare the responsibility chain node method and set the next responsibility chain node method.

[0167] Responsibility chain node interface: a standardized interface that defines how nodes handle tasks and the logic for maintaining the responsibility chain relationship.

[0168] Set next node method: used to dynamically bind the subsequent processing node of the current node.

[0169] For example: Define the interface ChainNode and declare the following methods: handle(Contextcontext): Executes the sensitive data detection task of the current node and modifies the context object.

[0170] setNext(ChainNodenextNode): Sets the next processing node of the current node.

[0171] Example interface code logic (specific code not shown): After processing each node, if there is a next node, call nextNode.handle(context) to pass the context object.

[0172] Step 5.2: Encapsulate the policy processor into a responsibility chain node and multiple responsibility chain nodes.

[0173] Policy processor: Implemented sensitive data identification methods (such as ID number detection class and phone number detection class).

[0174] Responsibility chain node: Wraps the policy processor as a chain node to support dynamic combination.

[0175] For example: Create a corresponding chain of responsibility node class for each policy processor, for example: IDCardNode: Inherits the ChainNode interface and calls the ID card number detection strategy.

[0176] PhoneNumberNode: inherits the ChainNode interface and calls the phone number detection strategy.

[0177] Inject the strategy instance into the node class (e.g. via constructor or dependency injection).

[0178] Step 5.3: Perform a specific type of sensitive data identification task in each responsibility chain node; the sensitive data identification task includes at least ID card number detection, phone number detection, and bank card number detection.

[0179] Specific type detection tasks: Each node focuses on a sensitive data type, such as: ID card number detection: Use regular expressions to match the 18-digit or 15-digit ID card format.

[0180] Phone number detection: matches 11 digits and common number segments (such as 13X, 18X).

[0181] Bank card number detection: Verify the card number length (16-19 digits).

[0182] For example: ID card number detection node: The regular expression matching strategy is called in the handle() method. If the match is successful, it is marked as sensitive in the context.

[0183] Phone number detection node: In the handle() method, the keyword library is traversed to check whether the field contains the mobile phone number keyword (such as "mobile phone" and "telephone").

[0184] Step 5.4: Encapsulate the intermediate results between the responsibility chain nodes into a context object and pass it to the next responsibility chain node for incremental detection.

[0185] Context object: A thread-safe data structure used to store the current detection status and intermediate results.

[0186] Incremental detection: Each node appends new detection results based on existing results in the context.

[0187] For example: Context object design: Contains fields: original data, current detection results (such as {ID number: sensitive, phone number: non-sensitive}), and confidence level.

[0188] Delivery logic: After the node detection is completed, the results are written to the context object (for example, context.setResult("ID number","sensitive")).

[0189] Call setNext() to continue processing the next node bound.

[0190] Step 5.5: Record the detection results of all responsibility chain nodes based on the preset concurrent hash table, aggregate multiple types of results to generate a comprehensive sensitive data report, and build a multi-type sensitive data detection task chain.

[0191] Concurrent hash table: A thread-safe hash table (such as ConcurrentHashMap) is used to aggregate multi-node detection results.

[0192] Comprehensive sensitive data report: Aggregates all node results and annotates the various sensitive statuses of fields.

[0193] For example: Result record: After each node is detected, the results are stored in a concurrent hash table with the field name + type as the key (such as user_phone: sensitive).

[0194] Report Generation: Traverse the hash table and count the number of sensitive types in each field.

[0195] Generate report example: Field name: user_info; Sensitive types: ID number (sensitive), phone number (sensitive), bank card number (non-sensitive); Comprehensive judgment: high sensitivity.

[0196] In a specific example, the E-government system needs to check whether the "Personal Data" field in the citizen information table contains the ID number, phone number, and bank card number. The system operates according to the following process: Building a chain of responsibility: Node sequence: IDCardNode—PhoneNumberNode—BankCardNode.

[0197] Passing a context object: The initial context contains the field content: "Zhang San, ID number 123456198001011234, phone number 13800138000".

[0198] IDCardNode detects the ID number and marks it as sensitive with a confidence level of 99%.

[0199] PhoneNumberNode detects the phone number and marks it as sensitive with a confidence level of 95%.

[0200] BankCardNode did not detect the bank card number and marked it as non-sensitive.

[0201] Generate a comprehensive report: The concurrent hash table records the results: {"ID number":"sensitive","phone number":"sensitive","bank card number":"non-sensitive"}.

[0202] The report shows that this field contains two types of sensitive data and requires priority processing.

[0203] Step 6: Perform sensitive data identification on the initial data set based on the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein the identification result includes the sensitive data type and probability.

[0204] Combining single-type parallel detection with multi-type chain of responsibility detection, sensitive data labeling results are generated and optimized in stages, and finally high-confidence results are determined through conflict verification to ensure the accuracy and comprehensiveness of the recognition results.

[0205] Step 6.1: Call the single-type detection execution template to perform parallel processing on the initial data set to generate the first sensitive data labeling result.

[0206] Single-type detection execution template: A predefined standardized detection process that supports multi-threaded parallel processing of a single type of sensitive data.

[0207] First sensitive data labeling result: preliminary detection result, including only a single sensitive type label and confidence level.

[0208] For example: Task allocation: Split the initial data set into multiple subtask blocks according to the thread pool configuration (for example, each block contains 1,000 data items).

[0209] Parallel detection: Each thread calls a single-type detection template and performs the following operations: Data preprocessing: Clean the data and unify the format (such as removing special characters).

[0210] Rule matching: Call the corresponding strategy based on the field type (such as regular expression matching ID number).

[0211] Result marking: The detection result (sensitive / non-sensitive), matching rule type and confidence level are stored in the intermediate cache.

[0212] Result aggregation: Merge the intermediate cache data of all threads to generate the first marked result set.

[0213] Step 6.2: Input the first sensitive data labeling result into the multi-type detection task chain, and perform detection in the responsibility chain order of the multi-type detection task chain to update the sensitive data type and probability of the first sensitive data labeling result to generate the second sensitive data labeling result.

[0214] Multi-type detection task chain: A detection process consisting of multiple responsibility chain nodes, each of which processes a specific sensitive data type.

[0215] Second sensitive data labeling result: The result after multi-type detection, including multiple sensitive type labels and updated confidence levels.

[0216] For example: Data input: Enter the first marking results into the responsibility chain one by one according to the data items.

[0217] Chain of responsibility testing process: Nodes are executed sequentially: the responsibility chain nodes are called in sequence (such as ID card number detection - phone number detection - bank card number detection).

[0218] Context propagation: Each node updates the sensitive type tag based on the context object (for example, if the phone number detection node finds new sensitive data, it appends the tag).

[0219] Confidence adjustment: Dynamically adjust the confidence level based on multi-node detection results (for example, if both the ID number and phone number are marked as sensitive, the confidence level is increased to 99%).

[0220] Generate second result: The updated context object is summarized as the second labeled result.

[0221] Step 6.3: Perform a conflict check on the second sensitive data marking result. If the same data entry is marked as sensitive by multiple responsibility chain nodes, the result with the highest confidence is taken as the recognition result.

[0222] Conflict check: This solves the problem where the same data entry is marked as sensitive by multiple nodes but with inconsistent confidence levels.

[0223] Confidence priority: The detection result with the highest confidence is used as the final judgment basis.

[0224] For example: Conflict detection: Traverse the second tagging results and filter out records where the same data entry has multiple sensitive type tags.

[0225] Confidence comparison: For example, if a data entry is marked as sensitive by the ID card number node (confidence level 90%) and is also marked as sensitive by the telephone number node (confidence level 85%), the ID card number detection result will be retained first.

[0226] Result solidification: Write the final marking results into the database or report file. The format example is as follows: Data ID: record_001; Field name: user_info; Sensitive type: ID number; Confidence level: 95% Conflict resolution: take the result with the highest confidence.

[0227] In a specific example, the F financial platform needs to identify sensitive data in the "Contact Information" field of the customer information table: Single type detection generates the first result: The mobile phone number detection template was called to process 100,000 pieces of data in parallel, marking 8,000 pieces of sensitive data containing mobile phone numbers with a confidence level of 80%-95%.

[0228] Update results of multi-type task chains: Responsibility chain order: mobile phone number detection - email detection - address keyword detection.

[0229] Context passing: A piece of data is marked as sensitive by the mobile phone number node (confidence level 90%), the email node does not detect sensitive information, and the address node is marked as sensitive (confidence level 70%).

[0230] Updated results: Marked as "Mobile number sensitive" (90% confidence) and "Address sensitive" (70% confidence).

[0231] Conflict check and final result: The data contains two sensitive type tags. After comparing the confidence levels, the mobile phone number detection result (with higher confidence) is retained.

[0232] The final report shows: Data ID: client_045; Field name: Contact information; Sensitive type: mobile phone number; Confidence level: 90%.

[0233] Step 7: Compare the identification results with the preset sensitive rule library to mark high-risk data and generate a visual report.

[0234] The risk level of the identification results is assessed through a predefined sensitive rule library, high-risk data is marked, and structured and visual reports are generated to provide decision support for subsequent risk disposal.

[0235] Step 7.1: Determine a sensitive rule base; wherein the sensitive rule base includes a predefined sensitive data type and risk level mapping table.

[0236] Sensitive rule base: A database or configuration file containing sensitive data types and their corresponding risk levels, used to define the risk levels of different sensitive data.

[0237] Risk level mapping table: This table clearly defines the correspondence between sensitive data types and risk levels (such as high, medium, and low).

[0238] For example: Rule base design: Create a configuration file (such as JSON or a database table) to store sensitive types and their risk levels.

[0239] Example mapping table: Sensitive type, risk level; ID number, high; Mobile phone number, Chinese; Bank card number, high; Home address, low;.

[0240] Dynamic loading rules: When the system starts, it reads the configuration file, loads the mapping table into the memory, and supports dynamic modification of the rule base content through the hot update mechanism.

[0241] Step 7.2: Query the risk level mapping table based on the sensitive type in the identification result to determine the risk level.

[0242] Risk level query: Based on the sensitive type in the identification result, the corresponding risk level is matched from the rule library.

[0243] For example: Traverse the recognition results: For each sensitive field of data (such as "ID number"), retrieve its risk level in the rule library.

[0244] Matching logic: If the sensitive type exists in the mapping table, the corresponding risk level is returned (for example, "ID Number" matches "High" risk).

[0245] If there is no match, it is marked as "unknown risk" and logged.

[0246] Step 7.3: If the risk level is higher than the preset threshold, the identification result will be marked as high-risk data. Structured report: Data report organized in a fixed format (such as JSON, CSV) to facilitate subsequent processing or import into other systems.

[0247] High-risk data: data entries whose risk level exceeds a preset threshold (such as "high" or confidence level ≥80%).

[0248] For example: Threshold configuration: Define the risk level threshold in the system profile (e.g. "High" or a numerical threshold of 80%).

[0249] Marking process: If the risk level of the data is "high", or the confidence level is ≥ the threshold, it is marked as high risk.

[0250] Example tagging results: Data ID: data_001; Field name: user_info; Sensitive type: ID number; Risk level: High; Confidence level: 95%.

[0251] Step 7.4: Integrate high-risk data and risk levels into a structured report.

[0252] Structured reports: Data reports organized in a fixed format (such as JSON, CSV), including a list of high-risk data and statistical information.

[0253] For example: Data Aggregation: High-risk data is classified and summarized by field name, sensitivity type, risk level, confidence level, etc.

[0254] Example report format: Data ID: data_001; Field name: user_info; Sensitive type: ID number; Risk level: High; Confidence level: 95 Statistics: Number of high-risk: 150; Number of medium risk: 300; Number of low risks: 50.

[0255] Step 7.5: Generate a visual report based on the structured report.

[0256] Visual report: Visually display risk distribution and sensitive type statistical results in the form of charts.

[0257] For example: Tool Selection: Integrate open source visualization libraries (such as ECharts, Matplotlib) or commercial tools (such as Tableau).

[0258] Chart content: Risk distribution pie chart: displays the proportion of high, medium, and low risk data.

[0259] Sensitive Type Histogram: Displays the number of detections and confidence distribution for each sensitive type.

[0260] Output format: Generate PDF, HTML, or interactive web reports for download or online viewing.

[0261] In a specific example, Bank I needs to perform a risk analysis on the "Remarks" field in customer transaction records: Determine the rule base: The mapping table defines "Bank Card Number" as high risk and "Mobile Number" as medium risk.

[0262] Query risk level: In the recognition results, "Bank Card Number" matches high risk (92% confidence), and "Mobile Number" matches medium risk (75% confidence).

[0263] Marking high-risk data: Only "Bank Card Number" is marked as high risk due to a confidence level ≥ 80%.

[0264] Generate a structured report: Summarize 200 high-risk data and 500 medium-risk data.

[0265] Visual display: The pie chart shows that high risk accounts for 28.6%, and the bar chart shows that "bank card number" has the largest number of detections.

[0266] Step 8: Adjust the data sampling strategy and sensitive data identification method through the hot update mechanism.

[0267] Use the hot update mechanism to dynamically load new sampling strategies or identification methods to replace or expand the existing strategy set, to a certain extent, ensure that this application continues to adapt to changes in business needs. The relevant content has been explained in step 1 and will not be repeated here.

[0268] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a sensitive data identification device based on big data, whose structure is as follows Figure 2 shown.

[0269] Figure 2 This is a schematic diagram of the internal structure of a sensitive data identification device based on big data provided by an embodiment of the present application. Figure 2 As shown, the equipment includes: at least one processor 201; and, a memory 202 communicatively coupled to the at least one processor; The memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein, the data sampling strategies include at least a full-volume extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; obtain metadata information of the target data source, and determine the data sampling strategy selection result based on the metadata information, and determine whether to start the data initialization operation; extract data from the target data source according to the selected data sampling strategy to form an initial data set; wherein, the data extraction strategies include a full-volume extraction strategy and a stratified sampling extraction strategy, and the initial data set includes the data items to be detected and the associated metadata; build a strategy factory, and match The target strategy in the data sampling strategy and the sensitive data identification method is matched to generate a single-type sensitive data detection execution template; wherein, the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; multiple policy processors are combined based on a preset chain of responsibility pattern to build a multi-type sensitive data detection task chain; sensitive data is identified on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein, the identification result includes the sensitive data type and probability; the identification result is compared with the preset sensitive rule library to mark high-risk data and generate a visual report; the data sampling strategy and sensitive data identification method are adjusted through a hot update mechanism.

[0270] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for sensitive data identification based on big data, storing computer executable instructions, wherein the computer executable instructions are set to: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein, the data sampling strategies include at least a full-volume extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; obtain metadata information of the target data source, and determine the data sampling strategy selection result based on the metadata information, and determine whether to start the data initialization operation; extract data from the target data source according to the selected data sampling strategy to form an initial data set; wherein, the data extraction strategies include a full-volume extraction strategy and a stratified sampling extraction strategy, and the initial data set includes the data items to be detected and the associated metadata; build a strategy factory, and match The target strategy in the data sampling strategy and the sensitive data identification method is matched to generate a single-type sensitive data detection execution template; wherein, the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; multiple policy processors are combined based on a preset chain of responsibility pattern to build a multi-type sensitive data detection task chain; sensitive data is identified on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein, the identification result includes the sensitive data type and probability; the identification result is compared with the preset sensitive rule library to mark high-risk data and generate a visual report; the data sampling strategy and sensitive data identification method are adjusted through a hot update mechanism.

[0271] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the IoT device and media embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.

[0272] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0273] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0274] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0275] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0276] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0277] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0278] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0279] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0280] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0281] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.

Claims

1. A sensitive data identification method based on big data, applied to a big data processing system, wherein the big data processing system includes multiple policy processors, characterized in that: The method comprises: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategies include at least a full data extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; Obtain metadata information of the target data source, determine a data sampling strategy selection result based on the metadata information, and determine whether to start a data initialization operation; Extracting data from the target data source according to a selected data sampling strategy to form an initial data set; wherein the data extraction strategy includes a full extraction strategy and a stratified sampling strategy; Constructing a policy factory and matching the data sampling strategy with the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template; wherein the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; Combine multiple policy processors based on a preset chain of responsibility model to build a multi-type sensitive data detection task chain; Performing sensitive data identification on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein the identification result includes the sensitive data type and probability; Compare the identification results with a preset sensitive rule library to mark high-risk data and generate a visual report; The data sampling strategy and sensitive data identification method are adjusted through the hot update mechanism.

2. The sensitive data identification method based on big data according to claim 1, characterized in that: Configuring data sampling strategies and sensitive data identification methods, and importing them into the big data processing system through a preset hot update mechanism, specifically including: Define the policy interface and declare data sampling strategies and sensitive data identification methods; Implementing a specific strategy class for the strategy interface; wherein the specific strategy class includes at least a full scan method, a stratified sampling method, a regular expression matching algorithm, and a keyword matching method; Encapsulate the specific policy class into an independent file, and configure an SPI service description file in the resource directory of the independent file to declare the policy implementation class contained in the independent file; Load the independent file based on the preset ServiceLoader to parse the SPI service description file and instantiate the policy implementation class; Injecting the instantiated policy implementation class into the memory of the big data processing system based on the preset Spring framework automatic assembly mechanism to form an extensible policy set; Listen to external configuration change requests, load new or updated independent files through a preset independent class loader, and replace the policy implementation class with the same name in the policy set to complete the hot update.

3. The sensitive data identification method based on big data according to claim 2, characterized in that: Obtaining metadata information of the target data source and determining whether to initiate a data initialization operation based on the metadata information specifically includes: Connecting to a target database of the target data source and extracting metadata information of the target database, wherein the metadata information includes field name, field type, field length, and index attributes; Determining whether the metadata information meets a preset sensitive data identification condition based on the field type and field length; wherein the sensitive data identification condition at least includes that the field length is greater than a preset minimum sensitive data length threshold; If the sensitive data identification condition is not met, the initialization operation of the metadata information is skipped; If the sensitive data identification condition is met, then calculating the data distribution characteristics of the metadata information, wherein the data distribution characteristics include data dispersion, data repetition rate and time window distribution; Determine a data sampling strategy selection result based on the data distribution characteristics, and generate an initialization operation start instruction; The metadata information, data sampling strategy selection result and initialization operation start instruction are stored in a preset configuration center.

4. The sensitive data identification method based on big data according to claim 1, characterized in that: Extract data from the target data source according to the selected data sampling strategy to form an initial data set, specifically including: If the full extraction strategy is selected, all data entries in the target data source are extracted to generate an initial data set; If a stratified sampling strategy is selected, the target data source is divided into multiple sub-data sources according to a time window or business rules, and the first number of data entries are extracted from each sub-data source according to a preset ratio; The extracted data entries are format-standardized to generate the initial dataset.

5. The sensitive data identification method based on big data according to claim 3 is characterized in that: A policy factory is built based on the policy pattern. The target policy is matched from the data sampling strategy and sensitive data identification method to generate a single-type sensitive data detection execution template, which specifically includes: Filtering a matching sensitive data identification method from the extensible policy set according to the field type; Create an execution template for the sensitive data identification method, defining the data preprocessing, rule matching, and result output processes; wherein the data preprocessing includes format conversion and null value checking; Associating the execution template with the thread pool, and adjusting the number of concurrent threads and task allocation strategy according to the initial data set size; Obtain key-value pair data and distribute it to multiple threads according to the preset concurrency number to perform single-type sensitive data detection in parallel; The detection results are marked as sensitive or non-sensitive, and the matching rule type and confidence are recorded in the intermediate result cache to build a single-type sensitive data detection execution template; wherein, the intermediate result cache adopts a concurrent data structure.

6. The sensitive data identification method based on big data according to claim 5, characterized in that: Combine multiple policy processors based on the preset chain of responsibility model to build a multi-type sensitive data detection task chain, including: Define the interface of the responsibility chain node, declare the responsibility chain node method and set the next responsibility chain node method; Encapsulating the policy processor as a chain of responsibility node, and multiple chain of responsibility nodes; Perform a specific type of sensitive data identification task at each responsibility chain node; wherein the sensitive data identification task includes at least identification card number detection, telephone number detection, and bank card number detection; Encapsulate the intermediate results between the responsibility chain nodes as a context object and pass it to the next responsibility chain node for incremental detection; Based on the preset concurrent hash table, the detection results of all responsibility chain nodes are recorded, and multi-type results are aggregated to generate a comprehensive sensitive data report to build a multi-type sensitive data detection task chain.

7. The sensitive data identification method based on big data according to claim 1, characterized in that: Performing sensitive data identification on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result specifically includes: Calling a single-type detection execution template to perform parallel processing on the initial data set to generate a first sensitive data labeling result; Inputting the first sensitive data labeling result into the multi-type detection task chain, and performing detection in the responsibility chain order of the multi-type detection task chain to update the sensitive data type and probability of the first sensitive data labeling result, so as to generate a second sensitive data labeling result; A conflict check is performed on the second sensitive data marking result. If the same data entry is marked as sensitive by multiple responsibility chain nodes, the result with the highest confidence is taken as the recognition result.

8. The sensitive data identification method based on big data according to claim 1, characterized in that: Compare the identification results with the preset sensitive rule library to mark high-risk data and generate a visual report, including: Determine a sensitive rule base; wherein the sensitive rule base includes a predefined sensitive data type and risk level mapping table; Querying the risk level mapping table according to the sensitive type in the identification result to determine the risk level; If the risk level is higher than a preset threshold, the identification result is marked as high-risk data; integrating the high-risk data and the risk levels into a structured report; A visual report is generated based on the structured report.

9. A sensitive data identification device based on big data, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategies include at least a full data extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; Obtain metadata information of the target data source, determine a data sampling strategy selection result based on the metadata information, and determine whether to start a data initialization operation; Extracting data from the target data source according to a selected data sampling strategy to form an initial data set; wherein the data extraction strategy includes a full extraction strategy and a stratified sampling strategy; Constructing a policy factory and matching the data sampling strategy with the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template; wherein the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; Combine multiple policy processors based on a preset chain of responsibility model to build a multi-type sensitive data detection task chain; Performing sensitive data identification on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein the identification result includes the sensitive data type and probability; Compare the identification results with a preset sensitive rule library to mark high-risk data and generate a visual report; The data sampling strategy and sensitive data identification method are adjusted through the hot update mechanism.

10. A non-volatile computer storage medium for sensitive data identification based on big data, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Configure data sampling strategies and sensitive data identification methods, and import them into the big data processing system through a preset hot update mechanism; wherein the data sampling strategies include at least a full data extraction strategy and a stratified sampling extraction strategy, and the sensitive data identification methods include single-type detection and multi-type detection; Obtain metadata information of the target data source, determine a data sampling strategy selection result based on the metadata information, and determine whether to start a data initialization operation; Extracting data from the target data source according to a selected data sampling strategy to form an initial data set; wherein the data extraction strategy includes a full extraction strategy and a stratified sampling strategy; Constructing a policy factory and matching the data sampling strategy with the target strategy in the sensitive data identification method to generate a single-type sensitive data detection execution template; wherein the single-type sensitive data detection execution template processes tasks in parallel through a thread pool; Combine multiple policy processors based on a preset chain of responsibility model to build a multi-type sensitive data detection task chain; Performing sensitive data identification on the initial data set according to the single-type detection execution template and the multi-type detection task chain to generate an identification result; wherein the identification result includes the sensitive data type and probability; Compare the identification results with a preset sensitive rule library to mark high-risk data and generate a visual report; The data sampling strategy and sensitive data identification method are adjusted through the hot update mechanism.

Citation Information

Cited By

  • Method, device and system for automatically identifying sensitive data to carry out de-identification processing

    CN121980617A