A processing method for preventing data leakage

By employing a layered matching mechanism and differentiated protection operations, the shortcomings of existing technologies in sensitive data identification and operation log monitoring are addressed. This enables refined data classification and real-time risk analysis, thereby improving the efficiency and security of data leakage protection.

CN121118115BActive Publication Date: 2026-02-03SICHUAN YOUJIA TRACEABILITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511657936.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-03
Estimated Expiration
2045-11-13

AI Technical Summary

Technical Problem

Existing data breach prevention technologies suffer from limitations such as limited methods for identifying sensitive data, difficulty in parsing the format and logical characteristics of structured data leading to missed detections or misjudgments, lack of refined classification mechanisms, and lack of real-time dynamic comparison capabilities in operation log monitoring, making it difficult to quickly identify potential leakage risks from abnormal access and unauthorized operations.

Method used

A layered matching mechanism is used to extract sensitive identification information from data files. Based on the sensitivity level, the data is divided into high, medium and low sensitivity data, and differentiated protection operations are performed for different levels, including hardware encryption, dynamic de-identification and role access control. Combined with a distributed log collection architecture that coordinates edge and cloud, real-time risk analysis is performed.

Benefits of technology

It enables comprehensive extraction of sensitive identifiers from both structured and unstructured data, avoiding the efficiency losses caused by homogeneous protection, improving the early warning and interception capabilities for data leakage risks, forming a fully intelligent protection system, and enhancing the accuracy and flexibility of data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121118115B_ABST
    Figure CN121118115B_ABST
Patent Text Reader

Abstract

The application discloses a processing method for preventing data leakage, and relates to the technical field of electric digital data processing, and comprises the following steps: obtaining data files, extracting sensitive identification information, dividing the data files into sensitive levels based on the sensitive identification information, obtaining sensitive level data, the sensitive level data comprising high sensitive data, medium sensitive data and low sensitive data, performing first, second and third protection operations on the high sensitive data, the medium sensitive data and the low sensitive data respectively, collecting operation logs of the data files in real time, extracting features of the operation logs, comparing the features of the operation logs with a preset condition library, obtaining a comparison result, extracting sensitive identification information from the data files by using a hierarchical matching mechanism, solving the problems of missed detection and misjudgment, dividing the data sensitive levels according to multi-dimensional grading, and differentiating protection, forming whole-process protection through an edge cloud collaborative architecture, and improving the accuracy, flexibility and real-time performance of data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a processing method for preventing data leakage. Background Technology

[0002] In recent years, with the acceleration of digitalization, the risk of data breaches has become increasingly severe. The security protection of core corporate data and sensitive user privacy information has become the focus of attention in various fields. From economic losses caused by the leakage of confidential information in business competition to the trust crisis caused by the illegal circulation of user privacy information, data breaches can have a serious impact on organizational operations, user rights and social order. Therefore, researching precise, dynamic and efficient technologies to prevent and handle data breaches has urgent practical significance.

[0003] Existing data breach prevention technologies generally have some shortcomings. Sensitive data identification methods are limited, with traditional methods relying solely on single text keyword matching, which is insufficient to effectively analyze the format and logical features in structured data, leading to missed detections or misjudgments of sensitive information. Protection strategies lack a refined hierarchical mechanism, employing homogeneous protection measures for data of different sensitivity levels. This can result in either excessive protection impacting business efficiency or insufficient protection leading to the exposure of high-risk data. Operation log monitoring remains at the basic recording level, lacking the ability to dynamically compare and identify anomalies in real time, making it difficult to quickly identify potential leakage risks from abnormal access and unauthorized operations. Summary of the Invention

[0004] The technical problem addressed by this invention is that existing data leakage prevention technologies generally have some shortcomings. Sensitive data identification methods are limited, and traditional methods rely solely on single text keyword matching, which is insufficient to effectively parse the format and logical features in structured data, leading to missed detections or misjudgments of sensitive information. Protection strategies lack a refined hierarchical mechanism, and the use of homogeneous protection measures for data of different sensitivity levels may either affect business efficiency due to over-protection or lead to the exposure of high-risk data due to insufficient protection. Operation log monitoring remains at the basic recording level, lacking the ability to dynamically compare and identify anomalies in operation behavior in real time, making it difficult to quickly identify potential leakage risks from abnormal access and unauthorized operations.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a processing method for preventing data leakage, comprising the following steps:

[0006] Step S1: Obtain a data file, extract the sensitive identification information of the data file, and classify the data file into sensitive levels based on the sensitive identification information to obtain sensitive level data;

[0007] The sensitivity level data includes highly sensitive data, moderately sensitive data, and low sensitive data;

[0008] Step S2: Perform the first protection operation, the second protection operation, and the third protection operation for the highly sensitive data, the moderately sensitive data, and the lowly sensitive data, respectively.

[0009] Step S3: Collect operation logs of data files in real time, extract features from the operation logs, compare them with the features of the operation logs using a preset condition library, and obtain the comparison results.

[0010] As a preferred embodiment of the processing method for preventing data leakage according to the present invention, the data file includes structured fields and text content;

[0011] The sensitive identification information is extracted from the data file using a hierarchical matching mechanism based on a preset feature library;

[0012] The feature library includes structured feature templates and unstructured feature templates;

[0013] The structured feature template is a feature set based on format logic. The feature set includes the combination logic of ID card number numbers, the specific length and check digit logic of bank card number, the prefix identifier of commercial contract number, and the combination logic of date and serial number.

[0014] The unstructured feature template is a preset set of sensitive keywords, which includes classified terms, core technology names, and exclusive terms related to personal privacy.

[0015] The hierarchical matching mechanism includes:

[0016] The structured fields are validated and matched bit by bit using the structured feature template to extract structured sensitive identifier information that conforms to the format logic;

[0017] The text content is searched using the unstructured feature templates to match and extract unstructured sensitive keywords.

[0018] The structured sensitive identifier information and unstructured sensitive keywords are summarized, and duplicate structured sensitive identifier information is removed to obtain the sensitive identifier information in the data file;

[0019] The sensitive identification information includes structured sensitive identification information and unstructured sensitive keywords.

[0020] As a preferred embodiment of the processing method for preventing data leakage described in this invention, the highly sensitive data, moderately sensitive data, and low-sensitivity data are obtained by classifying them according to the sensitivity identification information through a classification logic.

[0021] The classified information classification logic divides the data files into high-sensitivity data, medium-sensitivity data, and low-sensitivity data based on the importance score of the sensitive identification information, the quantitative value of the harm after leakage, the credibility level of the data source, the classified attributes of the purpose of use, and the restriction level of the dissemination scope.

[0022] The highly sensitive data corresponds to the top-secret level hazard standard, the medium-sensitive data corresponds to the confidential level hazard standard, and the low-sensitive data corresponds to the secret level hazard standard.

[0023] The quantification of the severity of the hazard after the leak includes extremely severe, severe, and minor levels;

[0024] The importance rating includes high, medium, and low ratings;

[0025] The credibility levels of the data sources include extremely high credibility, high credibility, medium credibility, and low credibility;

[0026] The classified attributes of the intended use include core classified information, general classified information, and non-classified information;

[0027] The levels of restriction on the scope of dissemination include strict restriction, limited dissemination, and public permission.

[0028] In a preferred embodiment of the processing method for preventing data leakage described in this invention, the highly sensitive data determination conditions include:

[0029] The degree of harm caused by the leakage of the sensitive identification information is quantified as an "extremely serious" level.

[0030] The sensitive identification information is judged to be highly sensitive if it has a high importance score, a medium or low credibility level of the data source, or is used for core confidential purposes.

[0031] When the above-mentioned sensitive data determination conditions are met, including:

[0032] The degree of harm caused by the leakage of the sensitive identification information is quantified as a severity level.

[0033] If the sensitive identification information is rated as medium in importance and the dissemination scope is restricted to limited dissemination or the purpose of use is general confidentiality, it is determined to be medium sensitive data.

[0034] When the following conditions for determining low-sensitivity data are met:

[0035] The degree of harm caused by the leakage of the sensitive identification information is quantified as mild, and the importance score is low.

[0036] Data is classified as low-sensitivity data when the data source has a trust level of extremely high or high, the purpose of use is non-confidential, and the dissemination scope is restricted to public access.

[0037] As a preferred embodiment of the data leakage prevention processing method described in this invention, performing a first protection operation on the highly sensitive data includes:

[0038] The highly sensitive data is fully encrypted using the national cryptographic algorithm and verified by binding a hardware dongle.

[0039] The national cryptographic algorithm adopts a dynamic key generation mechanism, which updates the encryption key based on the unique chip identifier built into the hardware dongle at regular intervals.

[0040] When the hardware dongle is physically connected to the terminal device, it allows decryption access to the encrypted highly sensitive data;

[0041] When the physical connection between the hardware dongle and the terminal device is detected to be disconnected, a temporary lock is triggered on the highly sensitive data, prohibiting the decryption access to the encrypted highly sensitive data;

[0042] Performing a second protection operation on the aforementioned sensitive data includes:

[0043] Dynamic desensitization is adopted, which is an operation of replacing sensitive fields in the data based on the visitor's permissions;

[0044] The dynamic desensitization process divides the visitor permissions into administrator level, operation level and query level by pre-setting the correspondence between visitor permissions and desensitization granularity.

[0045] The desensitization granularity corresponding to the visitor permissions is as follows: administrator level corresponds to replacing the first field, operation level corresponds to replacing the second field, and query level corresponds to replacing the third field. The dynamic desensitization processing operation is adjusted synchronously according to the real-time changes of the visitor permissions.

[0046] An immutable traceability identifier is added to the de-identified medium-sensitive data. The traceability identifier records the operation execution time, operator, and permission information of the dynamic de-identification process.

[0047] As a preferred embodiment of the data leakage prevention method described in this invention, performing a third protection operation on the low-sensitivity data includes:

[0048] Configure role-based access control lists;

[0049] The roles are user groups with low-sensitivity data access permissions;

[0050] The access control list pre-sets the binding relationship between low-sensitivity data and corresponding access roles;

[0051] The access control list introduces a threshold parameter for the access frequency of low-sensitivity data.

[0052] When any role accesses low-sensitivity data more frequently than the frequency threshold within the first time threshold, a temporary downgrade of permissions is triggered, and the role's access to low-sensitivity data is prohibited within the subsequent second time threshold. The access control list is also linked in real time to the modification operations of data files that are currently at the low-sensitivity level.

[0053] If the content of a data file currently classified as low-sensitivity data changes, resulting in a change in the sensitivity level of the data file, then the access permissions of the corresponding role associated with the data file are updated.

[0054] As a preferred embodiment of the processing method for preventing data leakage according to the present invention, the step of real-time collection of operation logs of data files and extraction of features from the operation logs includes:

[0055] The operation log includes the access IP, operation type, and data transmission path;

[0056] When collecting operation logs of data files in real time, a distributed log aggregation architecture that coordinates edge nodes and the cloud is adopted.

[0057] The edge nodes are used for localized log caching and feature extraction of operations completed within the third time threshold.

[0058] The features include extracting the operation initiation time, sensitivity level identification information of the data files involved in the operation, and operation instruction code from the operation log;

[0059] The cloud platform performs a comprehensive log correlation analysis on the operation logs, and simultaneously performs geolocation tracking and tagging of the accessing IPs and binds them to network operator information.

[0060] As a preferred embodiment of the processing method for preventing data leakage described in this invention, the preset condition library includes a baseline feature set, a whitelist strategy, and dynamic threshold parameters. The baseline feature set is constructed based on historical normal operation logs and includes the regular access period, average operation duration, and standard data transmission volume range corresponding to high-sensitivity data, medium-sensitivity data, and low-sensitivity data, respectively.

[0061] The whitelist policy includes a list of authorized access IPs and a list of operation types that are allowed to be performed based on sensitivity level data;

[0062] The dynamic threshold parameter is dynamically adjusted based on the real-time access frequency of the data file. If the access frequency of the data file within the fourth time threshold exceeds the historical average by a times, the comparison threshold for the corresponding operation behavior of the data file is lowered.

[0063] As a preferred embodiment of the processing method for preventing data leakage according to the present invention, the operation log features are compared using a two-layer matching mechanism based on the preset condition library.

[0064] The two-layer matching mechanism includes a first-layer matching mechanism and a second-layer matching mechanism;

[0065] The first-level matching mechanism is an exact match, which compares the access IP and operation type in the operation log with the list of authorized access IPs and the list of operation types allowed to be executed by sensitive level data in the whitelist of the preset condition library one by one;

[0066] When the access IP belongs to the whitelist of access IPs and the operation type matches the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as an exact match;

[0067] When the access IP is not in the whitelist of access IPs or the operation type does not conform to the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as a mismatch.

[0068] As a preferred embodiment of the data leakage prevention method described in this invention, the second-layer matching mechanism is fuzzy matching. For the non-matching items marked by the first-layer matching mechanism, the temporal and correlation features of the operation behavior are extracted, and the similarity score with the benchmark feature set is calculated using a cosine similarity algorithm to obtain the matching result, including:

[0069] When the similarity score is lower than the preset first proportion threshold, it is judged as a high-risk matching result;

[0070] When the similarity score is within a preset range, it is determined to be a medium-risk matching result;

[0071] When the similarity score is higher than the second ratio threshold, it is judged as a low-risk matching result;

[0072] The high-risk, medium-risk, and low-risk matching results are summarized with the exact matching items to form a comparison result.

[0073] The beneficial effects of this invention are as follows: A layered matching mechanism enables comprehensive extraction of the format logic features of structured fields and sensitive identification information of unstructured text in data files, solving the problem of missed detections and false judgments caused by traditional single-keyword matching. Based on multi-dimensional confidentiality classification logic, data is divided into high-sensitivity, medium-sensitivity, and low-sensitivity levels, and targeted differentiated protection operations such as hardware encryption binding, dynamic permission desensitization, and role-based access control are performed, avoiding efficiency losses and security vulnerabilities caused by homogeneous protection. Through a distributed log collection architecture that coordinates edge and cloud, combined with a two-layer mechanism of precise whitelist matching and fuzzy comparison of baseline features, real-time risk analysis and anomaly identification of operational behaviors are achieved, significantly improving the early warning and interception capabilities of data leakage risks. This forms a full-process intelligent protection system covering sensitive identification, graded protection, and dynamic monitoring, effectively enhancing the accuracy, flexibility, and real-time performance of data security protection. Attached Figure Description

[0074] Figure 1 This is a flowchart illustrating the steps of a method for preventing data leakage, provided as an embodiment of the present invention. Detailed Implementation

[0075] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0076] Example, refer to Figure 1 As an embodiment of the present invention, a processing method for preventing data leakage is provided, comprising the following steps:

[0077] Step S1: Obtain the data file, extract the sensitive identification information of the data file, classify the data file into sensitive levels based on the sensitive identification information, and obtain sensitive level data;

[0078] Sensitive data is categorized into highly sensitive, moderately sensitive, and low sensitive data.

[0079] Step S2: Perform the first protection operation, the second protection operation, and the third protection operation for highly sensitive data, moderately sensitive data, and low sensitive data, respectively;

[0080] Step S3: Collect operation logs from data files in real time, extract features from the operation logs, and compare them with the features of the operation logs using a preset condition library to obtain the comparison results.

[0081] In one embodiment, a data file is first acquired, and sensitive identification information of the data file is extracted. Based on the sensitive identification information, the data file is classified into sensitivity levels to obtain sensitive level data. Sensitive level data includes highly sensitive data, moderately sensitive data, and low sensitive data. A first protection operation, a second protection operation, and a third protection operation are performed for the highly sensitive data, moderately sensitive data, and low sensitive data, respectively. Operation logs of the data file are collected in real time, and features of the operation logs are extracted. The features of the operation logs are compared with a preset condition library to obtain the comparison results.

[0082] The data file includes structured fields and text content;

[0083] Sensitive identification information is extracted from data files using a hierarchical matching mechanism based on a pre-defined feature library;

[0084] The feature library includes structured feature templates and unstructured feature templates;

[0085] The structured feature template is a feature set based on format logic. The feature set includes the combination logic of ID card number numbers, the specific length and check digit logic of bank card number, the prefix identifier of commercial contract number, and the combination logic of date and serial number.

[0086] The unstructured feature template is a preset set of sensitive keywords, which includes classified terms, core technology names, and exclusive terms related to personal privacy.

[0087] The hierarchical matching mechanism includes:

[0088] By performing bit-by-bit verification and matching on structured fields using structured feature templates, structured sensitive identifier information that conforms to the format logic is extracted;

[0089] Full-text retrieval of text content is performed using unstructured feature templates to match and extract unstructured sensitive keywords;

[0090] The structured sensitive identifier information and unstructured sensitive keywords are summarized, and duplicate structured sensitive identifier information is removed to obtain the sensitive identifier information in the data file;

[0091] Sensitive identification information includes structured sensitive identification information and unstructured sensitive keywords.

[0092] In one embodiment, a hierarchical matching mechanism is used to extract sensitive identification information from data files containing both structured fields and text content by pre-setting a feature library that includes structured and unstructured feature templates. The structured feature templates are based on format logic, including 18-digit combination logic for ID card numbers (including address code, date of birth code, sequence code, and check digit), specific length (e.g., 16-19 digits) and check digit logic for bank card numbers (conforming to Luhn algorithm rules), and prefix-date-serial number combination logic for commercial contract numbers (e.g., prefix "HT-" + date "20240601" + serial number "001"). The unstructured feature templates include pre-defined sensitive terms (e.g., top secret and internal secret), core technology names (e.g., quantum communication encryption protocols and autonomous driving decision algorithms), and personal privacy-specific designations (e.g., medical record details and home address). The method first verifies and matches structured fields bit by bit using structured feature templates. For example, it verifies the ID number field in a user information table bit by bit according to the 18-bit encoding rule and extracts the numbers that meet the rule as structured sensitive identifier information. Then, it performs full-text search of the text content using unstructured feature templates. For example, when a quantum communication encryption protocol is matched in a technical document, it is extracted as an unstructured sensitive keyword. Finally, the two types of information are summarized and duplicate structured sensitive identifier information is removed (such as keeping only one instance when the same bank card number appears repeatedly in multiple fields), resulting in sensitive identifier information that includes both structured and unstructured sensitive keywords.

[0093] Highly sensitive data, moderately sensitive data, and lowly sensitive data are classified into three categories based on sensitivity identification information and a classification logic.

[0094] The classification logic is based on the importance of sensitive information, the quantification of the harm caused by leakage, the credibility level of the data source, the confidentiality attribute of the purpose of use, and the restriction level of the dissemination scope, to classify data files into high-sensitivity data, medium-sensitivity data, and low-sensitivity data.

[0095] Highly sensitive data corresponds to the top-secret level hazard standard, medium-sensitive data corresponds to the confidential level hazard standard, and low-sensitive data corresponds to the secret level hazard standard.

[0096] The quantification of the severity of the hazard after a leak includes three levels: extremely severe, severe, and minor.

[0097] Importance ratings include high, medium, and low ratings;

[0098] The credibility levels of data sources include extremely high credibility, high credibility, medium credibility, and low credibility;

[0099] The classification of the intended use includes core classified information, general classified information, and non-classified information;

[0100] The levels of restriction on the scope of dissemination include strict restriction, limited dissemination, and public permission.

[0101] When the criteria for determining highly sensitive data are met include:

[0102] The severity of harm caused by the leakage of sensitive identification information is quantified as "extremely serious".

[0103] Data is considered highly sensitive if it has a high importance score, a medium or low credibility level for the data source, or is used for core classified purposes.

[0104] When the criteria for determining sensitive data are met, the following conditions must be met:

[0105] The severity level of harm caused by the leakage of sensitive identification information is quantified as a serious level.

[0106] If the importance score of sensitive information is medium and the dissemination scope restriction level is limited dissemination or the purpose of use is general confidentiality, it is judged as medium sensitive data.

[0107] When the criteria for determining low-sensitivity data are met include:

[0108] The severity of harm caused by the leakage of sensitive identification information is rated as mild, and the importance score is low.

[0109] Data is classified as low-sensitivity data when the data source has a trust level of extremely high or high, the purpose of use is non-confidential, and the dissemination scope is restricted to public access.

[0110] In one embodiment, by integrating the importance score of sensitive identification information (high score, medium score, and low score), the quantitative value of the harm after disclosure (extremely serious level, serious level, and minor level), the trust level of the data source (extremely high trust, high trust, medium trust, and low trust), the confidentiality attribute of the purpose of use (core confidential, general confidential, and non-confidential), and the level of restriction on the scope of dissemination (strict restriction, limited dissemination, and public permission), data files are divided into highly sensitive data, moderately sensitive data, and lowly sensitive data. Highly sensitive data corresponds to the top secret level of harm standard, moderately sensitive data corresponds to the confidential level of harm standard, and lowly sensitive data corresponds to the secret level of harm standard.

[0111] When the criteria for determining highly sensitive data are met include:

[0112] The severity of harm caused by the leakage of sensitive identification information is quantified as "extremely serious".

[0113] Data is considered highly sensitive if it has a high importance score, a medium or low credibility level for the data source, or is used for core classified purposes.

[0114] When the criteria for determining sensitive data are met, the following conditions must be met:

[0115] The severity level of harm caused by the leakage of sensitive identification information is quantified as a serious level.

[0116] If the importance score of sensitive information is medium and the dissemination scope restriction level is limited dissemination or the purpose of use is general confidentiality, it is judged as medium sensitive data.

[0117] When the following conditions for determining low-sensitivity data are met simultaneously:

[0118] The severity of harm caused by the leakage of sensitive identification information is rated as mild, and the importance score is low.

[0119] Data is classified as low-sensitivity data when the source of the data is classified as extremely high or high, the purpose of use is non-confidential, and the scope of dissemination is restricted to public access.

[0120] Performing the first protection action on highly sensitive data includes:

[0121] The national cryptographic algorithm is used to encrypt all fields of highly sensitive data and a hardware dongle is used for verification.

[0122] The national cryptographic algorithm adopts a dynamic key generation mechanism, which updates the encryption key based on the unique chip identifier built into the hardware dongle at regular intervals.

[0123] When the hardware dongle is physically connected to the terminal device, it allows decryption access to the encrypted highly sensitive data;

[0124] When the physical connection between the hardware dongle and the terminal device is detected to be disconnected, a temporary lock is triggered on the highly sensitive data, prohibiting decryption access to the encrypted highly sensitive data.

[0125] Performing secondary protection measures on sensitive data includes:

[0126] Dynamic data masking is employed, which involves replacing sensitive fields in the data based on visitor permissions.

[0127] Dynamic desensitization processing divides visitor permissions into administrator level, operation level and query level by pre-setting the correspondence between visitor permissions and desensitization granularity;

[0128] The desensitization granularity corresponding to visitor permissions is as follows: administrator level corresponds to replacing the first field, operation level corresponds to replacing the second field, and query level corresponds to replacing the third field. Moreover, the dynamic desensitization processing is adjusted synchronously according to the real-time changes in visitor permissions.

[0129] Add an immutable traceability identifier to the de-identified medium-sensitive data. The traceability identifier records the execution time, operator, and permission information of the dynamic de-identification process.

[0130] Performing third-party protection operations on low-sensitivity data includes:

[0131] Configure role-based access control lists;

[0132] Roles are user groups with access to low-sensitivity data;

[0133] The access control list pre-defines the binding relationship between low-sensitivity data and corresponding access roles;

[0134] Access control lists introduce access frequency threshold parameters for low-sensitivity data.

[0135] When any role accesses low-sensitivity data more frequently than the frequency threshold within the first time threshold, a temporary downgrade of permissions is triggered, and the role's access to low-sensitivity data is prohibited within the subsequent second time threshold. The access control list is also linked in real time to the modification operations of the data files currently at the low-sensitivity level.

[0136] If the content of a data file currently classified as low-sensitivity data changes, resulting in a change in the sensitivity level of the data file, then the access permissions of the corresponding role associated with the data file will be updated.

[0137] In one embodiment, the high-sensitivity data protection (first protection operation) employs national cryptographic algorithms to encrypt all fields of the high-sensitivity data. This is combined with a hardware dongle verification mechanism to strengthen access control. The national cryptographic algorithm uses a dynamic key generation mechanism, based on the unique chip identifier embedded in the hardware dongle, to update the encryption key every 24 hours (based on key security and business convenience settings), ensuring key uniqueness and timeliness. Decryption access to encrypted data is only permitted when the hardware dongle is physically connected to the terminal device via a USB interface (e.g., core enterprise patent documents require the dongle to be inserted for decryption and viewing). If a physical connection loss is detected, a temporary lock is immediately triggered, prohibiting decryption operations. Simultaneously, a PLC algorithm is introduced and a self-locking function is configured. The triggering method includes: real-time detection of the physical connection status between the hardware dongle and the terminal device; when a physical connection is detected to be disconnected, a disconnection trigger signal is output; the PLC controller responds to the disconnection trigger signal by triggering the embedded counter and timer; after the counter is triggered, the count value is added to 1 to obtain the cumulative number of disconnections; after the timer is triggered, it starts timing to obtain the current timing value, with the timing unit being seconds; a preset hardware dongle physical connection disconnection allowable threshold is set as the counting threshold, such as 3 disconnections, which can be configured according to security requirements; and a preset RAM lock duration is set as the timing threshold, such as 24 hours, which can be configured according to security requirements.

[0138] The cumulative number of disconnections is compared with a counting threshold. When the cumulative number of disconnections is less than the counting threshold, the physical connection status continues to be monitored in real time. When a physical connection restoration is detected, a connection restoration trigger signal is output. The PLC controller responds to this signal by resetting the timer and keeping the PLC self-locking switch in an untriggered state. When the cumulative number of disconnections equals the counting threshold, a self-locking trigger signal is output. The PLC controller responds to this signal by triggering the self-locking switch in the PLC ladder diagram. After the self-locking switch is triggered, a RAM lock control signal is output. The PLC controller responds to this signal by immediately locking the RAM containing highly sensitive data, blocking any form of RAM access request, and simultaneously... The timer restarts the timer to obtain the locked time value. This locked time value is compared with the timer threshold. When the locked time value is less than the timer threshold, the RAM remains locked and the self-locking switch is triggered. When the locked time value equals the timer threshold, a self-locking reset signal is output. In response to the self-locking reset signal, the PLC controller outputs a RAM unlocking control signal. The PLC controller, in response to the RAM unlocking control signal, releases the lock on the stored RAM and simultaneously resets the counter and timer, restoring the normal connection detection logic between the hardware dongle and the terminal device. Through the above process, unauthorized access is blocked from the hardware layer, the cryptographic layer, and the storage access layer, ensuring the security protection of highly sensitive data in abnormal disconnection scenarios.

[0139] Sensitive data protection (secondary protection measure) implements dynamic data masking. Data fields are replaced differently based on visitor permissions (divided into administrator, operator, and query levels). A pre-defined relationship exists between permissions and masking granularity: administrators replace 10% of fields (the first field is set to 10%, hiding only a small amount of non-critical information, such as retaining the surname portion of a name); operators replace 50% of fields (the second field is set to 50%, masking core sensitive content, such as replacing the middle 8 digits of a bank card number with asterisks); and query users replace 80% of fields (the third field is set to 80%, significantly protecting privacy, such as retaining only the province information in an address). Furthermore, the masking operation is adjusted synchronously with real-time changes in permissions (e.g., automatically when an operator-level user is promoted to administrator level). (Reducing the field replacement ratio) The dynamic desensitization processing method includes configuring the PLC timer and the timer's clock crystal oscillator verification mechanism. The verification method includes: before the desensitization operation is executed, extracting the timestamp of the visitor's operation and outputting a time verification trigger signal. The PLC controller responds to the time verification trigger signal, triggering the embedded timer and the timer's clock crystal oscillator module. After the clock crystal oscillator module is triggered, it outputs the current crystal oscillator time signal. The timer synchronously records the crystal oscillator running time to obtain the crystal oscillator timing value. A preset time synchronization error threshold is set as the verification threshold. The time difference between the crystal oscillator timing value and the visitor's operation timestamp is calculated. The time difference is compared with the verification threshold. When the time difference is less than the verification threshold, the time synchronization is deemed qualified, and the desensitization permission is output. Upon receiving a signal, the PLC controller responds to the desensitization enable trigger signal, allowing the dynamic desensitization operation to proceed normally. When the time difference is greater than or equal to the verification threshold, it outputs a crystal oscillator control trigger signal. The PLC controller responds to this signal, initiating the dynamic crystal oscillator control logic. This logic corrects the time deviation by adjusting the crystal oscillator oscillation frequency. Simultaneously, a timer continuously accumulates the crystal oscillator control duration, obtaining a control timing value. This time value is compared to the control threshold. If the control timing value is less than the control threshold, crystal oscillator control continues, with real-time time difference comparisons, until the time difference is less than the verification threshold. Once this time difference is less than the verification threshold, a desensitization enable trigger signal is output. The PLC controller responds to this signal, allowing the dynamic desensitization operation to proceed. However, if the control timing value equals the control threshold, and time synchronization is still not achieved... If the step is successful, a desensitization blocking control signal is output. The PLC controller responds to this signal, blocking the start of the dynamic desensitization operation. Simultaneously, a time synchronization anomaly alarm signal is output, triggering the alarm module to execute alarm actions and record the time synchronization anomaly log. During the dynamic desensitization operation, according to the preset permission and desensitization granularity correspondence, administrators replace 10% of fields, operators replace 50% of fields, and queryers replace 80% of fields, completing the differentiated field replacement. The desensitization operation is synchronized with real-time changes in visitor permissions. An immutable blockchain traceability identifier is added to the desensitized data, recording the desensitization execution time, operator ID, and permission level. For example, the Human Resources Department's salary table displays 80% of the desensitized fields to query-level personnel.It also allows for the traceability of data masking operation records. Through the above process, it avoids the risk of masking rules becoming invalid or unauthorized masking bypasses due to time synchronization anomalies, achieving auditability of data usage and security verification in the time dimension.

[0140] Low-sensitivity data protection (third-party protection) involves building a role-based access control list (RBAC) to categorize users with low-sensitivity data access permissions into specific roles (e.g., regular employee query role and partner browsing role), with a pre-defined binding relationship between data and roles. The RBAC introduces an access frequency threshold parameter, defaulting to 50 accesses within one hour (the first time threshold is set to one hour to match regular business hours). When any role's access frequency exceeds 50 times, a temporary downgrade is triggered, prohibiting access for the following 30 minutes (the second time threshold is set to 30 minutes to balance risk control and recovery efficiency). (For example, if a role frequently downloads public product manuals exceeding the threshold within one hour, access is temporarily prohibited for half an hour). Simultaneously, the list is linked to data modification operations in real time. If low-sensitivity data's sensitivity level is upgraded to medium sensitivity due to content changes (e.g., adding customer contact information), the corresponding role's access permissions are automatically revoked, ensuring dynamic matching between permissions and data sensitivity levels.

[0141] Real-time acquisition of operation logs from data files, extraction of log features, including:

[0142] The operation log includes the access IP, operation type, and data transfer path;

[0143] When collecting operation logs of data files in real time, a distributed log aggregation architecture that coordinates edge nodes and the cloud is adopted.

[0144] Edge nodes are used for localized log caching and feature extraction of operations completed within a third time threshold;

[0145] Features include extracting the operation initiation time, sensitivity level identification information of the data files involved in the operation, and operation instruction codes from the operation log;

[0146] The cloud performs a comprehensive log correlation analysis on the operation logs, and also performs geolocation tracking and tagging of the access IPs and binds them to network operator information.

[0147] In one embodiment, the real-time collected data file operation logs include the access IP (e.g., 192.168.3.20), operation type (e.g., read, download, and modify), and data transmission path (e.g., from the department server to the employee terminal). The collection process employs a distributed log aggregation architecture that coordinates edge nodes and the cloud. Edge nodes must complete localized log caching and feature extraction of the operation behavior within 50ms (the third time threshold is set at 50ms, based on a balance between real-time monitoring and resource consumption). The extracted features include the operation initiation time (accurate to milliseconds, e.g., 2025-08-25). The operation involves sensitive level identification information of data files (such as high-sensitivity data, medium-sensitivity data, and low-sensitivity data) and operation behavior instruction codes (such as DOWNLOAD_001 representing a download operation and MODIFY_002 representing a modification operation). For example, after an edge node detects that a device with IP 10.0.5.8 is performing a medium-sensitivity data download operation, it completes log caching and extracts the features of 15:22:18.456, medium-sensitivity data, and DOWNLOAD_001 within 50ms. The cloud receives the log data uploaded by each edge node, performs correlation analysis on the entire domain log (such as identifying abnormal behavior of the same IP accessing multiple types of sensitive data across regions in a short period of time), and performs geographical location tracing and marking of the accessing IP (such as locating it to Shenzhen, Guangdong Province) and binds it to network operator information (such as China Telecom).

[0148] The preset condition library includes a baseline feature set, a whitelist strategy, and dynamic threshold parameters. The baseline feature set is built based on historical normal operation logs and includes the regular access period, average operation duration, and standard data transmission volume range corresponding to high-sensitive data, medium-sensitive data, and low-sensitive data, respectively.

[0149] The whitelist policy includes a list of authorized access IPs and a list of permitted operation types based on sensitivity level data;

[0150] The dynamic threshold parameter is dynamically adjusted based on the real-time access frequency of the data file. If the access frequency of the data file within the fourth time threshold exceeds the historical average by a times, the comparison threshold for the corresponding operation behavior of the data file is lowered.

[0151] In one embodiment, the preset condition library includes a baseline feature set, a whitelist policy, and dynamic threshold parameters. The baseline feature set is constructed based on six months of historical normal operation logs and is categorized by data sensitivity level for regular access characteristics. For example, high-sensitivity data corresponds to regular access times of 9:00-18:00 on weekdays, an average operation time of 30±10 minutes, and a standard data transfer volume of 1-5MB; medium-sensitivity data corresponds to regular access times of 7:00-22:00, an average operation time of 15±5 minutes, and a transfer volume of 5-20MB; and low-sensitivity data corresponds to accessibility all day, an average operation time of 5±3 minutes, and a transfer volume of 20-100MB. The whitelist policy includes a list of authorized access IPs (such as the enterprise intranet IP range 192.168.1.0 / 24) and... A list of permitted operation types for sensitive data (e.g., highly sensitive data is only allowed to be read and encrypted for download, while low-sensitivity data is allowed to be read, downloaded, and publicly shared); dynamic threshold parameters, with a default statistical period of 24 hours (the fourth time threshold is set to 24 hours to match the business day cycle). When the access frequency of a data file exceeds twice the historical average (a is set to twice, based on a reasonable threshold of ±100% of historical data fluctuations), the comparison threshold for the corresponding operation behavior of that data is automatically lowered from 90% to 80% (for example, if 90% matching degree is required to determine normality, it is lowered to 80% when abnormal to improve detection sensitivity). For example, if the access volume of a low-sensitivity product manual suddenly increases by more than twice the historical average within 24 hours, the normal matching threshold for its download operation is automatically lowered to accurately identify high-frequency abnormal access.

[0152] The operation log features are compared using a two-layer matching mechanism based on the preset condition library.

[0153] The two-layer matching mechanism includes a first-layer matching mechanism and a second-layer matching mechanism;

[0154] The first-level matching mechanism is an exact match, which compares the access IP and operation type in the operation log with the list of authorized access IPs and the list of operation types allowed to be executed by sensitive level data in the whitelist of the preset condition library one by one;

[0155] When the access IP belongs to the whitelist of access IPs and the operation type matches the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as an exact match;

[0156] When the access IP is not in the whitelist of access IPs or the operation type does not conform to the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as a mismatch.

[0157] The second-layer matching mechanism is fuzzy matching. For the non-matching items marked by the first-layer matching mechanism, the temporal and association features of the operation behavior are extracted. The similarity score with the baseline feature set is calculated using the cosine similarity algorithm to obtain the matching result, including:

[0158] When the similarity score is lower than the preset first proportion threshold, it is judged as a high-risk matching result;

[0159] When the similarity score is within a preset range, it is determined to be a medium-risk matching result;

[0160] When the similarity score is higher than the second ratio threshold, it is judged as a low-risk matching result;

[0161] The high-risk, medium-risk, and low-risk matching results are summarized with the exact matching items to form a comparison result.

[0162] In one embodiment, the two-layer matching mechanism includes two layers of processing logic: a first layer of precise matching and a second layer of fuzzy matching. The first layer of precise matching compares the access IP (e.g., 192.168.1.101) and operation type (e.g., encrypted download) in the operation log with the authorized IP list (e.g., the enterprise intranet segment 192.168.1.0 / 24) and the list of allowed operation types for the corresponding sensitive data (e.g., highly sensitive data is only allowed to be read and encrypted download) in the preset condition database whitelist. If the IP is in the whitelist and the operation type matches the list (e.g., an intranet IP performs encrypted download of highly sensitive data), it is marked as a precise match and allowed directly. If the IP is not in the whitelist or the operation type does not match (e.g., a public IP performs ordinary download of highly sensitive data), it is marked as a mismatch and enters the second layer of fuzzy matching processing. The second layer of fuzzy matching extracts the temporal features (e.g., access time 22:30) and related features (e.g., data transfer volume 10MB, operation duration 45 minutes) of the operation behavior for the mismatches in the first layer. It then calculates the similarity score with the baseline feature set (constructed based on historical normal behavior, such as the usual access time for highly sensitive data 9:00-18:00, and the transfer volume 1-5MB) using the cosine similarity algorithm: a score below 60% (the first proportion threshold is set at 60%, based on the minimum matching degree of historical abnormal behavior) is judged as high risk (e.g., large-scale transfer of highly sensitive data during non-working hours in the early morning), 60%-85% (the preset proportion range is set at 60%-85%, covering the fluctuation range of normal behavior) is judged as medium risk (e.g., the transfer volume slightly exceeds the standard range during the afternoon access time), and above 85% (the second proportion threshold is set at 85%, matching the high similarity of normal behavior) is judged as low risk (e.g., normal download operation of low-sensitivity data). Finally, the exact match items are combined with the high-risk, medium-risk, and low-risk match results to form a comparison result (e.g., if a public IP is downloading sensitive data, and the fuzzy match score is 75%, it is judged as medium risk and the audit process is triggered).

[0163] This invention achieves comprehensive extraction of the format logic features of structured fields and sensitive identification information of unstructured text in data files through a layered matching mechanism. It solves the problem of missed detections and false judgments caused by traditional single keyword matching. Based on multi-dimensional confidentiality classification logic, it divides data into high-sensitivity, medium-sensitivity, and low-sensitivity levels, and performs differentiated protection operations such as hardware encryption binding, dynamic permission desensitization, and role-based access control. This avoids the efficiency loss and security vulnerabilities caused by homogeneous protection. Through a distributed log collection architecture that coordinates edge and cloud, combined with a two-layer mechanism of whitelist precise matching and baseline feature fuzzy comparison, it realizes real-time risk analysis and anomaly identification of operation behavior, significantly improving the early warning and interception capabilities of data leakage risks. It forms a full-process intelligent protection system covering sensitive identification, hierarchical protection, and dynamic monitoring, effectively enhancing the accuracy, flexibility, and real-time performance of data security protection.

[0164] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0165] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A processing method for preventing data leakage, characterized in that, Includes the following steps: Step S1: Obtain a data file, extract the sensitive identification information of the data file, and classify the data file into sensitive levels based on the sensitive identification information to obtain sensitive level data; The sensitivity level data includes highly sensitive data, moderately sensitive data, and low sensitive data; Step S2: Perform the first protection operation, the second protection operation, and the third protection operation for the highly sensitive data, the moderately sensitive data, and the lowly sensitive data, respectively. Step S3: Collect operation logs of data files in real time, extract features of the operation logs, compare them with the features of the operation logs using a preset condition library, and obtain the comparison results; The data file includes structured fields and text content; The sensitive identification information is extracted from the data file using a hierarchical matching mechanism based on a preset feature library; The feature library includes structured feature templates and unstructured feature templates; The structured feature template is a feature set based on format logic. The feature set includes the combination logic of ID card number numbers, the specific length and check digit logic of bank card number, the prefix identifier of commercial contract number, and the combination logic of date and serial number. The unstructured feature template is a preset set of sensitive keywords, which includes classified terms, core technology names, and exclusive terms related to personal privacy. The hierarchical matching mechanism includes: The structured fields are validated and matched bit by bit using the structured feature template to extract structured sensitive identifier information that conforms to the format logic; The text content is searched using the unstructured feature templates to match and extract unstructured sensitive keywords. The structured sensitive identifier information and unstructured sensitive keywords are summarized, and duplicate structured sensitive identifier information is removed to obtain the sensitive identifier information in the data file; The sensitive identification information includes structured sensitive identification information and unstructured sensitive keywords; The highly sensitive data, moderately sensitive data, and low sensitive data are classified according to the sensitive identification information through a classification logic. The classified information classification logic divides the data files into high-sensitivity data, medium-sensitivity data, and low-sensitivity data based on the importance score of the sensitive identification information, the quantitative value of the harm after leakage, the credibility level of the data source, the classified attributes of the purpose of use, and the restriction level of the dissemination scope. The highly sensitive data corresponds to the top-secret level hazard standard, the medium-sensitive data corresponds to the confidential level hazard standard, and the low-sensitive data corresponds to the secret level hazard standard. The quantification of the severity of the hazard after the leak includes extremely severe, severe, and minor levels; The importance rating includes high, medium, and low ratings; The credibility levels of the data sources include extremely high credibility, high credibility, medium credibility, and low credibility; The classified attributes of the intended use include core classified information, general classified information, and non-classified information; The levels of restriction on the scope of dissemination include strict restriction, limited dissemination, and public permission; Performing the first protection operation on the highly sensitive data includes: The highly sensitive data is fully encrypted using the national cryptographic algorithm and verified by binding a hardware dongle. The national cryptographic algorithm adopts a dynamic key generation mechanism, which updates the encryption key based on the unique chip identifier built into the hardware dongle at regular intervals. When the hardware dongle is physically connected to the terminal device, it allows decryption access to the encrypted highly sensitive data; When the physical connection between the hardware dongle and the terminal device is detected to be disconnected, a temporary lock is triggered on the highly sensitive data, prohibiting the decryption access to the encrypted highly sensitive data; Performing a second protection operation on the aforementioned sensitive data includes: Dynamic desensitization is adopted, which is an operation of replacing sensitive fields in the data based on the visitor's permissions; The dynamic desensitization process divides the visitor permissions into administrator level, operation level and query level by pre-setting the correspondence between visitor permissions and desensitization granularity. The desensitization granularity corresponding to the visitor permissions is as follows: administrator level corresponds to replacing the first field, operation level corresponds to replacing the second field, and query level corresponds to replacing the third field. The dynamic desensitization processing operation is adjusted synchronously according to the real-time changes of the visitor permissions. An immutable traceability identifier is added to the de-identified moderately sensitive data. The traceability identifier records the operation execution time, operator, and permission information of the dynamic de-identification process. Performing third-party protection on the low-sensitivity data includes setting up a role-based access control list.

2. The data leakage prevention method as described in claim 1, characterized in that: When the above-mentioned highly sensitive data determination conditions are met, including: The degree of harm caused by the leakage of the sensitive identification information is quantified as an "extremely serious" level. The sensitive identification information is judged to be highly sensitive if it has a high importance score, a medium or low credibility level of the data source, or is used for core confidential purposes. When the above-mentioned sensitive data determination conditions are met, including: The degree of harm caused by the leakage of the sensitive identification information is quantified as a severity level. If the sensitive identification information is rated as medium in importance and the dissemination scope is restricted to limited dissemination or the purpose of use is general confidentiality, it is determined to be medium sensitive data. When the following conditions for determining low-sensitivity data are met: The degree of harm caused by the leakage of the sensitive identification information is quantified as mild, and the importance score is low. Data is classified as low-sensitivity data when the data source has a trust level of extremely high or high, the purpose of use is non-confidential, and the dissemination scope is restricted to public access.

3. The data leakage prevention method as described in claim 2, characterized in that: Performing third protection operations on the aforementioned low-sensitivity data also includes: The roles are user groups with low-sensitivity data access permissions; The access control list pre-sets the binding relationship between low-sensitivity data and corresponding access roles; The access control list introduces a threshold parameter for the access frequency of low-sensitivity data. When any role accesses low-sensitivity data more frequently than the frequency threshold within the first time threshold, a temporary downgrade of permissions is triggered, and the role's access to low-sensitivity data is prohibited within the subsequent second time threshold. The access control list is also linked in real time to the modification operations of data files that are currently at the low-sensitivity level. If the content of a data file currently classified as low-sensitivity data changes, resulting in a change in the sensitivity level of the data file, then the access permissions of the corresponding role associated with the data file are updated.

4. The data leakage prevention method as described in claim 3, characterized in that: The operation log of the real-time acquired data file is used to extract features from the operation log, including: The operation log includes the access IP, operation type, and data transmission path; When collecting operation logs of data files in real time, a distributed log aggregation architecture that coordinates edge nodes and the cloud is adopted. The edge nodes are used for localized log caching and feature extraction of operations completed within the third time threshold. The features include extracting the operation initiation time, sensitivity level identification information of the data files involved in the operation, and operation instruction code from the operation log; The cloud platform performs a comprehensive log correlation analysis on the operation logs, and simultaneously performs geolocation tracking and tagging of the accessing IPs and binds them to network operator information.

5. The data leakage prevention method as described in claim 4, characterized in that: The preset condition library includes a baseline feature set, a whitelist strategy, and dynamic threshold parameters. The baseline feature set is constructed based on historical normal operation logs and includes the regular access period, average operation duration, and standard data transmission volume range corresponding to high-sensitive data, medium-sensitive data, and low-sensitive data, respectively. The whitelist policy includes a list of authorized access IPs and a list of operation types that are allowed to be performed based on sensitivity level data; The dynamic threshold parameter is dynamically adjusted based on the real-time access frequency of the data file. If the access frequency of the data file within the fourth time threshold exceeds the historical average by a times, the comparison threshold for the corresponding operation behavior of the data file is lowered.

6. The data leakage prevention method as described in claim 5, characterized in that: The operation log features are compared using a two-layer matching mechanism based on the preset condition library. The two-layer matching mechanism includes a first-layer matching mechanism and a second-layer matching mechanism; The first-level matching mechanism is an exact match, which compares the access IP and operation type in the operation log with the list of authorized access IPs and the list of operation types allowed to be executed by sensitive level data in the whitelist of the preset condition library one by one; When the access IP belongs to the whitelist of access IPs and the operation type matches the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as an exact match; When the access IP is not in the whitelist of access IPs or the operation type does not conform to the list of operation types allowed to be executed according to the corresponding sensitivity level data, it is marked as a mismatch.

7. The data leakage prevention method as described in claim 6, characterized in that: The second-layer matching mechanism is fuzzy matching. For the non-matching items marked by the first-layer matching mechanism, the temporal and association features of the operation behavior are extracted. The similarity score with the baseline feature set is calculated using the cosine similarity algorithm to obtain the matching result, including: When the similarity score is lower than the preset first proportion threshold, it is judged as a high-risk matching result; When the similarity score is within a preset range, it is determined to be a medium-risk matching result; When the similarity score is higher than the second ratio threshold, it is judged as a low-risk matching result; The high-risk, medium-risk, and low-risk matching results are summarized with the exact matching items to form a comparison result.

Citation Information

Patent Citations

  • A document desensitization system and method based on big data

    CN109284631A

  • Data security management platform for preventing data loss

    CN115733681A