Data leakage prevention method and device

By acquiring and processing initial document data, generating key data and generating target document data, the data leakage problem caused by key loss or encrypted data corruption in the prior art is solved, and precise data protection and access control are achieved.

CN119989426AActive Publication Date: 2025-05-13BEIJING GUODU INTERNET TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510476032.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In existing data leakage prevention technology, when the key is lost or encrypted data is corrupted, the data may be inaccessible, resulting in data leakage.

Method used

By obtaining initial document data, obtaining sensitive information based on fingerprint information, generating key data, and generating target document data based on initial document data, key data and operation data of sensitive information, to achieve accurate data protection and access control.

Benefits of technology

Ensure that only true sensitive information is protected, avoid omissions or misjudgment, and achieve accurate identification and protection of sensitive information, preventing unauthorized access and data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989426A_ABST
    Figure CN119989426A_ABST
Patent Text Reader

Abstract

The invention provides a data leakage prevention method and device, which are applied to the technical field of data processing, and the method comprises the following steps: obtaining initial document data; obtaining sensitive information in the initial document data based on the fingerprint information; generating key data corresponding to the initial document data based on the sensitive information in the initial document data and the access permission of the initial document data; and generating target document data based on the initial document data, the key data corresponding to the initial document data and the operation data of the sensitive information. According to the method, the potential leakage risk can be found and prevented in time, and the security and privacy of the data can be effectively protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing technology, and in particular to a data leakage prevention method and device. Background Art

[0002] With the rapid development of information technology, data has become the core asset of enterprises and organizations. However, data leakage incidents occur frequently, causing huge losses to enterprises and organizations, including customer loss, loss of credibility, loss of core technology, legal issues and economic compensation. Therefore, data leakage prevention technology has become an important means to ensure data security.

[0003] At present, data leakage prevention generally uses encryption algorithms to encrypt data to ensure the security of data during storage and transmission. Common encryption technologies include disk encryption, file encryption, and transparent document encryption and decryption.

[0004] However, in the above method, once the key is lost or the encrypted data is damaged, the data may be unable to be recovered, thereby causing data leakage. Summary of the invention

[0005] Based on the above problems, the present invention is proposed to provide a data leakage prevention method and device that overcomes the above problems or at least partially solves the above problems.

[0006] According to one aspect of the present invention, a data leakage prevention method is provided, comprising the following steps: Get initial document data; Acquire sensitive information in the initial document data based on the fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of a sensitive level; Based on the sensitive information in the initial document data and the access rights to the initial document data, generate key data corresponding to the initial document data; Target document data is generated based on the initial document data, key data corresponding to the initial document data, and operation data of sensitive information.

[0007] Optionally, based on the sensitive information in the initial document data and the access rights to the initial document data, key data corresponding to the initial document data is generated, including: Obtain keywords from sensitive information in the initial document data; Calculate the contribution of keywords in the sensitive information in the initial document data; Get the access frequency of the initial document data; Key data corresponding to the initial document data is generated based on keywords in the sensitive information in the initial document data, contribution of the keywords in the sensitive information in the initial document data, access rights of the initial document data, and access frequency of the initial document data.

[0008] Optionally, calculating the contribution of keywords in the sensitive information in the initial document data includes: Calculate the product of the frequency square term of the keyword in the sensitive information in the initial document data and the access data volume of the keyword in the sensitive information in the initial document data to obtain a first product; Calculate the product of the score of the keyword in the sensitive information in the initial document data and the length influence factor of the keyword in the sensitive information in the initial document data to obtain a second product; Calculate the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data; Perform word segmentation on the initial document data and count the total number of all words in the initial document data to obtain the total number of words in the initial document data; The contribution of the keyword in the sensitive information in the initial document data is determined based on the first product, the second product, the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data, and the total number of words in the initial document data.

[0009] Optionally, determining the contribution of the keywords in the sensitive information in the initial document data based on the first product, the second product, the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data, and the total number of words in the initial document data includes: Add the first product to the second product to obtain a first sum; Subtract the first sum from the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data to obtain a first difference value; Obtain a first preset value and a second preset value; Calculating a ratio of the total number of words in the initial document data to a first preset value to obtain a first ratio; Calculate the sum of the first ratio and the second preset value to obtain a second sum value; The ratio of the first difference value to the second sum value is calculated to obtain the contribution of the keyword in the sensitive information in the initial document data.

[0010] Optionally, based on keywords in the sensitive information in the initial document data, contribution of the keywords in the sensitive information in the initial document data, access rights to the initial document data, and access frequency of the initial document data, key data corresponding to the initial document data is generated, including: According to the contribution of the keyword in the sensitive information in the initial document data and the logarithmic transformation of the keyword in the sensitive information in the initial document data, a corresponding keyword score is configured for the keyword in the sensitive information in the initial document data; Configure corresponding permission scores for the access rights of the initial document data according to the permission level; Obtain a frequency score based on the proportion of the access frequency of the initial document data to all document data; Obtain the total word count score based on the proportion of the total word count of the initial document data in all document data; Obtain the criticality scores of keywords in the sensitive information in the initial document data through keyword scores, permission scores, frequency scores, and total word count scores; The keywords in the sensitive information in the initial document data are sorted in descending order according to the criticality scores of the keywords in the sensitive information in the initial document data, and the key data corresponding to the initial document data are output.

[0011] Optionally, the method further comprises: Calculate the length impact factor of the keyword in the sensitive information in the initial document data; Calculate the length impact factor of the keywords in the sensitive information in the initial document data, including: Acquire the length and third preset value of the keyword in the sensitive information in the initial document data; Calculate the difference between the third preset value and the length of the keyword in the sensitive information in the initial document data to obtain a second difference value; The ratio of the length of the keyword in the sensitive information in the initial document data to the second difference is calculated to obtain a length influence factor of the keyword in the sensitive information in the initial document data.

[0012] Optionally, generating target document data based on the initial document data, key data corresponding to the initial document data, and operation data of sensitive information includes: According to the key data corresponding to the initial document data, access type and index information of the key data corresponding to the initial document data are obtained; Obtain target operation items in the operation data of sensitive information; Target document data is generated based on the initial document data, target operation items in the sensitive information operation data, key data corresponding to the initial document data, and access type and index information of the key data corresponding to the initial document data.

[0013] Optionally, generating target document data based on the initial document data, the target operation item in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data includes: Obtaining the position of the key data in the initial document data through the index position in the index information of the key data corresponding to the initial document data; By defining a dictionary to configure the target operation item in the operation data of sensitive information and the sensitive information processing rules corresponding to the target operation item; Detect the key data at the location of the key data in the initial document data according to the access type of the key data corresponding to the initial document data, and obtain sensitive information in the key data; Obtaining a target operation item in the operation data of sensitive information in the key data, and processing the sensitive information using a sensitive information processing rule corresponding to the target operation item to obtain processed key data; The key data corresponding to the initial document data is updated with the processed key data to generate target document data.

[0014] Optionally, obtaining sensitive information in the initial document data based on the fingerprint information includes: Obtain user ID through fingerprint information; The initial document data is traversed by user ID, and a preset function is used to search for sensitive information in the initial document data.

[0015] According to another aspect of the present invention, there is provided a data leakage prevention device, comprising: A data acquisition module, used to acquire initial document data; An information extraction module, used for obtaining sensitive information in the initial document data based on the fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of a sensitive level; A key data generation module, used to generate key data corresponding to the initial document data based on sensitive information in the initial document data and access rights to the initial document data; The target document generation module is used to generate target document data based on the initial document data, key data corresponding to the initial document data, and operation data of sensitive information.

[0016] According to the solution of the present invention, in the present invention, the initial document data is first obtained, and then the sensitive information in the initial document data is obtained based on the fingerprint information. By obtaining the sensitive information in the initial document data based on the fingerprint information, the key data in the document can be accurately located and identified, ensuring that only the real sensitive information is protected to avoid omissions or misjudgments. Then, based on the sensitive information in the initial document data and the access rights of the initial document data, the key data corresponding to the initial document data is generated. The key data is generated based on the access rights of the initial document data to ensure that only authorized users can access the sensitive information. This access control mechanism can effectively prevent unauthorized access and data leakage. Thus, based on the initial document data, the key data corresponding to the initial document data and the operation data of the sensitive information, the target document data is generated. Through the operation data of the sensitive information, the access and operation behavior of the sensitive information can be monitored in real time, and the potential leakage risks can be discovered and prevented in time, which can effectively protect the security and privacy of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flow chart of a data leakage prevention method according to an embodiment of the present invention is shown; Figure 2 A structural block diagram of a data leakage prevention device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0018] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0019] like Figure 1 As shown, a data leakage prevention method proposed in this embodiment includes the following steps: Step S101: Acquire initial document data.

[0020] Initial document data refers to document data stored outside the enterprise information system in personal computers, file servers, mailboxes, WeChat, QQ and other media, including word, excel, ppt, pdf, etc. These document data may contain sensitive information of the enterprise, such as design documents, design drawings, source code, marketing plans, financial statements and other content involving state secrets and corporate trade secrets.

[0021] In data leakage prevention management, the initial document data needs to be classified and graded. Generally, it is divided into 3-5 levels, and different protection strategies are set for documents of different levels. For example, sensitive-level document data is prohibited from being sent out, while public-level document data is not subject to any management measures.

[0022] The initial document data is generally composed of a leadership group, an operations group, and a business interface group. The leadership group is responsible for the overall promotion of data leakage prevention management affairs, the operations group is responsible for the overall planning, coordination, and operation of data leakage prevention work, and the business interface person is the main force in promoting the implementation of data leakage prevention.

[0023] Step S102: Acquire sensitive information in the initial document data based on the fingerprint information.

[0024] In order to obtain sensitive information in the initial document data based on the fingerprint information, it is necessary to first obtain the user ID through the fingerprint information, then traverse the initial document data through the user ID, and use a preset function to search for sensitive information in the initial document data.

[0025] Specifically, when a user visits a website, the front end obtains the browser fingerprint through JavaScript and generates a user ID, and then sends the user ID to the server for storage. The server obtains all the document data of the user (i.e., the initial document data) based on the user ID. The client traverses the initial document data and uses the preset sensitive information detection function to find sensitive information.

[0026] Assume that a company needs to detect sensitive information in employees' document data to prevent the leakage of sensitive content such as business secrets and customer information. The company uses a document management system in which employees' document data is stored. At the same time, the company deploys a sensitive information detection system to identify sensitive information in documents.

[0027] The implementation steps are as follows: The client obtains document data User login: Employees log in to the document management system through the client. After the system verifies the user's identity, it allows the user to access document data within his or her authority.

[0028] Document list acquisition: The client sends a request to the document management system to obtain a list of all documents of the user. The document management system retrieves all documents of the user from the database or file storage based on the user ID and returns the document list to the client.

[0029] Client traverses document data Document list display: After the client receives the document list, it will be displayed to the user on the interface. The user can see all his documents, including document name, creation time, modification time and other information.

[0030] Traversing documents: The client starts traversing each document in the document list one by one. For each document, the client requests the document management system to obtain the detailed content of the document. The document management system reads the document content from the storage based on the document ID and returns it to the client.

[0031] Use the preset sensitive information detection function to find sensitive information Sensitive information detection function: The client has a built-in preset sensitive information detection function, which identifies sensitive information based on a series of predefined rules. These rules may include keyword matching, regular expressions, data format checks, etc. For example, the detection function will check whether the document contains keywords such as "confidential", "financial data", "customer information", "account password", etc., or check whether the document contains data that conforms to the format of ID card number, bank card number, email address, etc.

[0032] Document content analysis: The client passes the content of each document to the sensitive information detection function. The detection function analyzes the document content line by line or paragraph by paragraph, and checks whether there is sensitive information according to the preset rules.

[0033] Sensitive information marking: If the detection function finds sensitive information in the document, it will mark the sensitive information and record relevant information, such as the document name, sensitive information content, location, etc. The client stores the marked sensitive information in a separate list or report for subsequent processing.

[0034] Sensitive information handling and feedback Sensitive information report generation: The client generates a sensitive information report based on the detection results. The report lists in detail all documents containing sensitive information and their related information. The report can be presented to the user in a table, list, or other visual form.

[0035] User feedback and processing: The client displays the sensitive information report to the user, who can review the report content and confirm whether there are any false positives. If the user confirms that the information in the document is indeed sensitive, the client provides options for the user to further process the document, such as encrypting, deleting, or marking it as processed. At the same time, the client can also send the sensitive information report to the company's security administrator for subsequent security audits and processing.

[0036] By traversing the document data and using the preset sensitive information detection function, the enterprise can timely discover sensitive information in employee documents and effectively reduce the risk of data leakage. At the same time, the process can be automated to improve work efficiency and reduce the workload of manual review.

[0037] Step S103: Generate key data corresponding to the initial document data based on the sensitive information in the initial document data and the access rights of the initial document data.

[0038] In order to generate key data corresponding to the initial document data based on the sensitive information in the initial document data and the access rights of the initial document data, it is necessary to first obtain the keywords in the sensitive information in the initial document data, then calculate the contribution of the keywords in the sensitive information in the initial document data, and then obtain the access frequency of the initial document data, so as to generate the key data corresponding to the initial document data based on the keywords in the sensitive information in the initial document data, the contribution of the keywords in the sensitive information in the initial document data, the access rights of the initial document data and the access frequency of the initial document data.

[0039] The embodiments of the present application can more accurately identify which documents are critical data by comprehensively considering the keywords of sensitive information, the contribution of keywords, access frequency and access rights, avoiding misjudgment caused by relying solely on a single factor. In addition, resource allocation can be optimized to help enterprises focus limited security resources on critical data and improve the efficiency and effectiveness of data protection. For highly critical documents, more stringent security measures can be taken, such as encryption, access control, auditing, etc.; for low-criticality documents, the intensity of security measures can be appropriately reduced to save resources.

[0040] Specifically, in order to obtain the keywords in the sensitive information in the initial document data, sensitive information detection must first be performed. Sensitive information detection can use a preset sensitive information detection tool (such as based on keyword matching, regular expressions or machine learning models) to scan the document data and identify the sensitive information in the document.

[0041] Then, keywords are extracted, that is, keywords are extracted from the identified sensitive information. For example, if the document contains sensitive information such as "financial statements", "customer lists", and "contract terms", these words are extracted as keywords.

[0042] Calculating the contribution of keywords in sensitive information in initial document data includes: first calculating the product of the frequency square term of the keywords in the sensitive information in the initial document data and the access data volume of the keywords in the sensitive information in the initial document data to obtain a first product, then calculating the product of the score of the keywords in the sensitive information in the initial document data and the length influence factor of the keywords in the sensitive information in the initial document data to obtain a second product, and then calculating the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data, and then performing word segmentation processing on the initial document data and counting the total number of all words in the initial document data to obtain the total number of words in the initial document data, thereby determining the contribution of the keywords in the sensitive information in the initial document data based on the first product, the second product, the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data and the total number of words in the initial document data.

[0043] Among them, based on the first product, the second product, the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data and the total number of words in the initial document data, the contribution of the keywords in the sensitive information in the initial document data is determined, including: adding the first product and the second product to obtain a first sum; subtracting the first sum from the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data to obtain a first difference; obtaining a first preset value and a second preset value; calculating the ratio of the total number of words in the initial document data to the first preset value to obtain a first ratio; calculating the sum of the first ratio and the second preset value to obtain a second sum; calculating the ratio of the first difference to the second sum to obtain the contribution of the keywords in the sensitive information in the initial document data.

[0044] The formula for calculating the contribution of keywords in the sensitive information in the initial document data is as follows: , in, Indicates keywords in sensitive information in the initial document data. Indicates the frequency of keywords in the sensitive information in the initial document data, Indicates the amount of access data for keywords in sensitive information in the initial document data. represents the score of the keyword in the sensitive information in the initial document data, Indicates the length impact factor of the keywords in the sensitive information in the initial document data, Represents the total number of words in the initial document data, represents the first preset value, Indicates the second preset value.

[0045] The embodiment of the present application integrates multiple dimensions such as keyword frequency, access data volume, keyword score, length impact factor, total number of words in the document, etc., and can comprehensively evaluate the contribution of keywords to document sensitivity. Compared with the evaluation method of a single dimension, this comprehensive evaluation method can more accurately reflect the importance and sensitivity of keywords in the document, and avoid misjudgment due to ignoring certain key factors. In addition, by accurately calculating the contribution of each keyword, the truly critical sensitive information in the document can be quickly located. For example, in a document containing a large amount of ordinary information and a small amount of sensitive information, this method can accurately identify those keywords that play a decisive role in the sensitivity of the document, such as "financial statements" and "customer lists", thereby helping enterprises or organizations focus their attention and resources on these key information and improve the efficiency of data management and protection.

[0046] Furthermore, the introduction of the access data volume in the embodiment of the present application enables the calculation of the keyword contribution to dynamically reflect the real-time status of the document. The higher the access frequency of the document, the more important it is in actual use, and it may involve more business activities and interactions with sensitive information. By incorporating the access data volume into the calculation, keywords that are frequently accessed and contain sensitive information in actual use can be discovered in a timely manner, thereby better coping with the dynamically changing data environment. As the document content is updated and the access situation changes, the contribution of the keyword will also be adjusted accordingly. This dynamic adaptability enables enterprises or organizations to continuously and accurately grasp the changes in the sensitivity of the document, take corresponding data protection measures in a timely manner, and ensure that key data is always under effective control. Furthermore, by calculating the contribution of keywords in the embodiment of the present application, enterprises can clearly identify which keywords contribute the most to the sensitivity of the document, thereby focusing limited security resources and management efforts on these key data. Based on the accurately calculated keyword contribution, enterprises can protect key data more specifically and prevent sensitive information leakage. By identifying and protecting those keywords that contribute the most to the sensitivity of the document, the risk of data leakage can be effectively reduced, and important assets such as the company's business secrets and customer information can be protected.

[0047] Among them, calculating the length influence factor of the keyword in the sensitive information in the initial document data includes: obtaining the length of the keyword in the sensitive information in the initial document data and a third preset value; calculating the difference between the third preset value and the length of the keyword in the sensitive information in the initial document data to obtain a second difference value; calculating the ratio of the length of the keyword in the sensitive information in the initial document data to the second difference value to obtain the length influence factor of the keyword in the sensitive information in the initial document data.

[0048] Specifically, the formula for calculating the length impact factor of the keyword in the sensitive information in the initial document data is as follows: , in, Indicates the length of the keyword in the sensitive information in the initial document data. Indicates the third preset value.

[0049] The embodiment of the present application takes into account the impact of keyword length on sensitivity, and by calculating the length impact factor, it is possible to balance the impact of keyword length on its sensitivity, making the calculation of contribution more reasonable and accurate. For example, "finance" and "financial statements" are both related to finance, but "financial statements" are more specific, and their length impact factor will reflect this difference, thereby giving a more reasonable weight in the contribution calculation, while avoiding length deviation. If the impact of keyword length is not considered, it may lead to misjudgment of sensitivity. For example, a shorter keyword may appear frequently in a document, but if its length is shorter, its actual sensitivity may not be high. By introducing the length impact factor, the sensitivity assessment bias caused by differences in keyword length can be avoided, and the reliability of the assessment results can be improved.

[0050] Since the distribution of keyword lengths may be different in different languages ​​and document types. By adjusting the third preset value, the method can adapt to multiple languages ​​and document types, and enhance the flexibility and adaptability of the evaluation. For example, in a Chinese document, the keyword may be shorter; while in an English document, the keyword may be longer. By reasonably setting the third preset value, it can be ensured that the length impact factor can effectively reflect the sensitivity of the keyword in different languages ​​and document types. In addition, the calculation formula of the length impact factor quantifies the relationship between the keyword length and the preset value, providing a scientific quantitative indicator for sensitivity assessment. This quantitative method makes the evaluation process more objective and operational, and avoids the uncertainty caused by subjective judgment. By calculating the length impact factor, the contribution of the keyword length to its sensitivity can be accurately measured, thereby giving a more accurate weight in the contribution calculation.

[0051] Suppose an enterprise needs to conduct a sensitive information assessment on documents in its internal document management system to determine which keywords contribute most to the sensitivity of the documents. These documents may contain sensitive content such as business secrets, customer information, financial data, etc. The enterprise hopes to calculate the contribution of each keyword through scientific methods in order to better manage and protect these documents.

[0052] The implementation steps are as follows: 1. Calculate the product of the frequency square of the keyword and the amount of access data (first product) Keyword frequency statistics: Analyze each document and count the frequency of each sensitive keyword in the document. For example, the keyword "financial report" appears 5 times in a document.

[0053] Access data acquisition: Get the access data of the document from the access log of the document management system, that is, the total number of times the document has been accessed. Assume that the document has been accessed 100 times in the past month.

[0054] Calculate the frequency square term: square the frequency of the keyword. For example, the frequency square term of the keyword "financial statements" is 5^2 = 25.

[0055] Calculate the first product: multiply the frequency squared term by the amount of accessed data. For example, the first product is 25×100 =2500.

[0056] 2. Calculate the product of the keyword score and the length impact factor (second product) Keyword score setting: A score is pre-set based on the sensitivity of the keyword. For example, the score for "financial statement" is 10.

[0057] Keyword length statistics: Calculate the length of the keyword, for example, "financial report" has 4 characters. Third preset value setting: Set a third preset value, assuming it is 10. Calculate length impact factor: Calculate the difference between the third preset value and the keyword length (second difference): 10 - 4 = 6.

[0058] Calculate the ratio of the keyword length to the second difference (length impact factor): 4 / 6 = 0.67. Calculate the second product: multiply the keyword score by the length impact factor. For example, the second product is 10×0.67 = 6.7.

[0059] 3. Calculate the logarithmic transformation of the access data volume of keywords Logarithmic transformation: Perform a logarithmic transformation on the access data volume of the keyword. Assuming the natural logarithm is used, the logarithm of 100 is 4.6.

[0060] 4. Segment the document and count the total number of words Word segmentation: Perform word segmentation on the document to split the document content into individual words. For example, if the document content is "This is a detailed description of the financial statements", after word segmentation, you will get "This is a detailed description of the financial statements".

[0061] Count the total number of words: Count the total number of words after word segmentation. For example, the total number of words in this document is 8.

[0062] 5. Determine the contribution of keywords To calculate the first sum: add the first product to the second product. For example, the first sum is 2500 + 6.7 = 2506.7.

[0063] Calculate the first difference: Subtract the first sum from the logarithmic transformation of the access data volume. For example, the first difference is 2506.7 - 4.6 = 2502.1.

[0064] Preset value setting: Set the first preset value and the second preset value. Suppose the first preset value is 100 and the second preset value is 5.

[0065] Calculate the first ratio: calculate the ratio of the total number of words to the first preset value. For example, the first ratio is 8 / 100=0.08.

[0066] Calculate the second sum: add the first ratio to the second preset value. For example, the second sum is 0.08 + 5 = 5.08.

[0067] Calculate the contribution: Calculate the ratio of the first difference to the second sum. For example, the contribution is 2502.1 / 5.08=492.54.

[0068] The embodiment of the present application can more accurately evaluate the contribution of each keyword to the sensitivity of the document by comprehensively considering multiple factors such as keyword frequency, access data volume, keyword length, etc. In addition, the contribution of keywords can be dynamically adjusted according to the access frequency and content changes of the document to ensure that the evaluation results always reflect the current status of the document. Furthermore, it helps enterprises identify the keywords that contribute most to the sensitivity of the document, so as to focus resources on these key points, improve the efficiency of data protection, better protect key data, prevent the leakage of sensitive information, and improve overall data security.

[0069] Among them, based on the keywords in the sensitive information in the initial document data, the contribution of the keywords in the sensitive information in the initial document data, the access rights of the initial document data and the access frequency of the initial document data, the key data corresponding to the initial document data is generated, including: configuring corresponding keyword scores for the keywords in the sensitive information in the initial document data according to the contribution of the keywords in the sensitive information in the initial document data and the logarithmic transformation of the keywords in the sensitive information in the initial document data; configuring corresponding permission scores for the access rights of the initial document data according to the permission level; obtaining a frequency score according to the proportion of the access frequency of the initial document data in all document data; obtaining a total word count score according to the proportion of the total number of words in the initial document data in all document data; obtaining a criticality score for the keywords in the sensitive information in the initial document data through the keyword score, permission score, frequency score and total word count score; sorting the keywords in the sensitive information in the initial document data in descending order according to the criticality scores of the keywords in the sensitive information in the initial document data, and outputting the key data corresponding to the initial document data.

[0070] Based on the above implementation steps, we first configure the keyword score by taking a logarithmic transformation of the keyword contribution 492.54, assuming it is 6.2. Then, we configure the keyword score based on the logarithmic transformation value of the keyword contribution 6.2. Assuming that the logarithmic transformation value is directly used as the keyword score, the score of the keyword "financial statement" is about 6.2.

[0071] Configure permission score: The access permission of the document is "Senior Management" and the preset permission level is "High". Configure the permission score according to the permission level. Assume that the "High" permission score is 8.

[0072] Get frequency score: The access frequency of the document is 100 times, the total access frequency of all documents is 1000 times, and the frequency share = 0.1. Configure the frequency score based on the frequency share. Assuming that the frequency share is multiplied by 100 as the frequency score, the frequency score is 10.

[0073] Get the total word count score: The total word count of the document is 8, and the total word count of all documents is 1000, so the total word count ratio = 0.008. Configure the total word count score based on the total word count ratio. Assuming that the total word count ratio is multiplied by 100 as the total word count score, the total word count score is 0.8.

[0074] Calculate the keyword’s criticality score: keyword score = 6.2, authority score = 8, frequency score = 10, total word score = 0.8, criticality score = 25.

[0075] Sort the criticality scores of all keywords in descending order: Suppose another keyword "customer list" has a criticality score of 20, and another keyword "meeting minutes" has a criticality score of 5.

[0076] The sorting results are: financial statements (25), customer lists (20), meeting minutes (5).

[0077] Output key data: Output the sorted keywords and their key scores as the key data of the document.

[0078] The embodiment of the present application can accurately identify the truly critical sensitive information in the document by comprehensively considering factors such as keyword contribution, access rights, access frequency, and the total number of words in the document. This method avoids misjudgment caused by relying solely on a single factor and improves the accuracy and reliability of the evaluation. In addition, the calculation of keyword contribution takes into account the dynamic changes in access frequency and the total number of words in the document, and can promptly reflect the actual use and sensitivity changes of the document. This dynamic adaptability enables enterprises to continuously and accurately grasp the sensitivity status of documents and adjust data protection strategies in a timely manner.

[0079] Furthermore, by calculating the key scores of keywords and sorting them, enterprises can clearly identify which keywords contribute most to the sensitivity of the document, thereby focusing limited security resources on these key data. This method improves resource utilization efficiency and ensures that key data is protected. This method can effectively identify and protect key data and prevent sensitive information leakage. By identifying and protecting those keywords that contribute most to the sensitivity of documents, enterprises can reduce the risk of data leakage and protect important assets such as business secrets and customer information.

[0080] Step S104: Generate target document data based on the initial document data, key data corresponding to the initial document data, and operation data of sensitive information.

[0081] In order to generate target document data based on the initial document data, the key data corresponding to the initial document data, and the operation data of sensitive information, it is necessary to first obtain the access type and index information of the key data corresponding to the initial document data based on the key data corresponding to the initial document data, and then obtain the target operation item in the operation data of the sensitive information, and then generate the target document data based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data.

[0082] The target operation items include at least access control operations, anomaly detection operations, and backup operations.

[0083] There is an internal document management system in the enterprise, which stores various sensitive and important document data, such as employees' personal information, financial reports, business secrets, etc.

[0084] The initial document data takes a document containing employee personal information as an example, in which key data includes the employee’s ID number, bank card number, home address, etc.

[0085] Get the access type and index information of the key data corresponding to the initial document data, where access type: through system permission settings, determine that only specific personnel in the human resources department (such as personnel specialists) and employees themselves have access rights. For example, personnel specialists can view and edit these key data, and employees themselves can only view their own information. Index information: Create indexes for these key data, such as indexing by fields such as employee number and name, for quick positioning and retrieval.

[0086] The target operation items in the operation data of obtaining sensitive information are as follows: Access control operation: record each access behavior to key data, including access time, visitor identity, access purpose, etc. For example, when a human resources specialist inquires about an employee's ID number, the system will record the operation. Abnormal detection operation: the system will monitor the operation behavior on key data in real time. If an abnormal situation is found, such as frequent access to the key data of the same employee, the visitor identity does not match the operation authority, etc., an alarm will be issued immediately. For example, if an employee outside the human resources department attempts to access the bank card number of another employee, the system will detect this abnormal behavior and record it. Backup operation: regularly back up key data to ensure data security and recoverability. For example, every night the system will automatically back up the key data of all employees and store the backup files on a secure server.

[0087] The target document data is generated based on the initial document data, the target operation items in the sensitive information operation data, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data. The target document data not only contains the original employee personal information, but also includes the following: Access records: record every access to key data, including access time, visitor identity, access purpose, etc.

[0088] Anomaly detection records: records abnormal behaviors detected by the system, including specific descriptions of the abnormal behaviors, occurrence time, key data involved, etc.

[0089] Backup records: record the time of the backup operation, the storage location of the backup file, the integrity check of the backup file, and other information.

[0090] The embodiments of the present application ensure that only authorized personnel can access key data and prevent data leakage through access control operations. Abnormal detection operations can promptly detect and prevent potential security threats and protect key data from malicious attacks. In addition, index information makes the retrieval of key data more efficient and can quickly locate and obtain required information. At the same time, the backup operation ensures the recoverability of data. Even if data is lost or damaged, the data can be restored through backup files.

[0091] Furthermore, the target document data contains access records, anomaly detection records, and backup records, which provide strong support for data auditing and tracing. Enterprises can view the access and operation history of key data at any time to ensure the compliance of data use, and quickly locate the cause and take measures when problems arise.

[0092] Wherein, based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data, the target document data is generated, including: obtaining the position of the key data in the initial document data through the index position in the index information of the key data corresponding to the initial document data; configuring the target operation item in the operation data of the sensitive information and the sensitive information processing rule corresponding to the target operation item by defining a dictionary; performing detection on the key data at the position of the key data in the initial document data through the access type of the key data corresponding to the initial document data to obtain the sensitive information in the key data; obtaining the target operation item in the operation data of the sensitive information in the key data, and processing the sensitive information using the sensitive information processing rule corresponding to the target operation item to obtain the processed key data; updating the key data corresponding to the initial document data through the processed key data to generate the target document data.

[0093] Suppose a company has a customer information management system that stores detailed customer information, including name, ID number, bank card number, contact number, home address, etc. Now it is necessary to manage and protect this sensitive information to ensure data security and compliance.

[0094] The initial document data is a table containing customer information, including: customer number, customer name, ID number, bank card number, contact number, home address, etc.

[0095] The system sets index information for each key data field to facilitate quick location and access. The index information is as follows: ID number: Column 3 Bank card number: Column 4 The system defines a dictionary of sensitive information operation data, which contains the target operation items and the corresponding sensitive information processing rules: 1. Access Control: Rule: Only authorized personnel are allowed to access sensitive information. Action: Log the identity of the visitor and the time of access.

[0096] 2. Data desensitization: Rules: Desensitize sensitive information and hide some information. Operation: Hide the middle 8 digits of the ID card number and the middle 10 digits of the bank card number.

[0097] 3. Anomaly Detection: Rules: Detect abnormal access behaviors, such as frequent access to the same sensitive information. Actions: Log abnormal behaviors and issue alerts.

[0098] Based on the index information, the system determines the location of the key data in the initial document data: The ID number is in the 3rd column.

[0099] The bank card number is in the 4th column.

[0100] Depending on the access type of the key data (e.g., "access control"), the system performs detection operations at the location of the key data. For example: It is detected that user A has accessed customer 001's ID number and bank card number.

[0101] The system records the visitor as user A and the access time as 14:00 on March 26, 2025.

[0102] The system processes the detected sensitive information according to the defined sensitive information processing rules: Process the ID number "123456789012345678" into "1234XXXXXXXX5678" (hide the middle 8 digits).

[0103] Process the bank card number "62284800000000000000" into "622848XXXXXXXXXX000" (hide the middle 10 digits).

[0104] The system updates the processed key data into the initial document data to generate the final target document data.

[0105] The embodiments of the present application ensure that only authorized personnel can access sensitive information and prevent data leakage by recording the visitor identity and access time. Abnormal access behavior can be discovered and blocked in a timely manner through anomaly detection, and sensitive information can be protected from malicious attacks. Sensitive information can also be desensitized and some information can be hidden. Even if the data is leaked, the complete sensitive information cannot be directly obtained, thereby protecting user privacy.

[0106] In addition, the embodiments of the present application also quickly locate the location of key data through index information, thereby improving data retrieval efficiency, and realize automated processing by defining sensitive information operation data and processing rules, thereby reducing manual intervention and improving processing efficiency.

[0107] Furthermore, through strict access control, data desensitization and anomaly detection measures, it is ensured that data processing complies with the requirements of relevant laws and regulations, and legal risks caused by data leakage and other issues are avoided. Through the above embodiments, it can be seen that this data processing method can effectively improve data security and management efficiency, protect user privacy, and meet regulatory requirements. It is applicable to customer information management systems within enterprises, customer information management systems of financial institutions, and citizen information management systems of government departments.

[0108] According to the solution of the present invention, in the present invention, the initial document data is first obtained, and then the sensitive information in the initial document data is obtained based on the fingerprint information. By obtaining the sensitive information in the initial document data based on the fingerprint information, the key data in the document can be accurately located and identified, ensuring that only the real sensitive information is protected to avoid omissions or misjudgments. Then, based on the sensitive information in the initial document data and the access rights of the initial document data, the key data corresponding to the initial document data is generated. The key data is generated based on the access rights of the initial document data to ensure that only authorized users can access the sensitive information. This access control mechanism can effectively prevent unauthorized access and data leakage. Thus, based on the initial document data, the key data corresponding to the initial document data and the operation data of the sensitive information, the target document data is generated. Through the operation data of the sensitive information, the access and operation behavior of the sensitive information can be monitored in real time, and the potential leakage risks can be discovered and prevented in time, which can effectively protect the security and privacy of the data.

[0109] Figure 2 FIG. 4 shows a data leakage prevention device proposed by an embodiment of the present invention. Figure 2 As shown, the device comprises: Data acquisition module 201, used to acquire initial document data; The information extraction module 202 is used to obtain sensitive information in the initial document data based on the fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of a sensitive level; A key data generating module 203, configured to generate key data corresponding to the initial document data based on sensitive information in the initial document data and access rights of the initial document data; The target document generation module 204 is used to generate target document data based on the initial document data, key data corresponding to the initial document data, and operation data of sensitive information.

[0110] Optionally, the key data generation module 203 is further used to obtain keywords in the sensitive information in the initial document data; Calculate the contribution of keywords in the sensitive information in the initial document data; Get the access frequency of the initial document data; Key data corresponding to the initial document data is generated based on keywords in the sensitive information in the initial document data, contribution of the keywords in the sensitive information in the initial document data, access rights of the initial document data, and access frequency of the initial document data.

[0111] Optionally, the key data generating module 203 is further used to calculate the product of the frequency square term of the keyword in the sensitive information in the initial document data and the access data volume of the keyword in the sensitive information in the initial document data to obtain a first product; Calculate the product of the score of the keyword in the sensitive information in the initial document data and the length influence factor of the keyword in the sensitive information in the initial document data to obtain a second product; Calculate the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data; Perform word segmentation on the initial document data and count the total number of all words in the initial document data to obtain the total number of words in the initial document data; The contribution of the keyword in the sensitive information in the initial document data is determined based on the first product, the second product, the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data, and the total number of words in the initial document data.

[0112] Optionally, the key data generating module 203 is further configured to add the first product and the second product to obtain a first sum value; Subtract the first sum from the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data to obtain a first difference value; Obtain a first preset value and a second preset value; Calculating a ratio of the total number of words in the initial document data to a first preset value to obtain a first ratio; Calculate the sum of the first ratio and the second preset value to obtain a second sum value; The ratio of the first difference value to the second sum value is calculated to obtain the contribution of the keyword in the sensitive information in the initial document data.

[0113] Optionally, the key data generating module 203 is further configured to configure corresponding keyword scores for the keywords in the sensitive information in the initial document data according to the contribution of the keywords in the sensitive information in the initial document data and the logarithmic transformation of the keywords in the sensitive information in the initial document data; Configure corresponding permission scores for the access rights of the initial document data according to the permission level; Obtain a frequency score based on the proportion of the access frequency of the initial document data to all document data; Obtain the total word count score based on the proportion of the total word count of the initial document data in all document data; Obtain the criticality scores of keywords in the sensitive information in the initial document data through keyword scores, permission scores, frequency scores, and total word count scores; The keywords in the sensitive information in the initial document data are sorted in descending order according to the criticality scores of the keywords in the sensitive information in the initial document data, and the key data corresponding to the initial document data are output.

[0114] Optionally, the apparatus further comprises: calculating a length impact factor of a keyword in the sensitive information in the initial document data; Calculate the length impact factor of the keywords in the sensitive information in the initial document data, including: Acquire the length and third preset value of the keyword in the sensitive information in the initial document data; Calculate the difference between the third preset value and the length of the keyword in the sensitive information in the initial document data to obtain a second difference value; The ratio of the length of the keyword in the sensitive information in the initial document data to the second difference is calculated to obtain a length influence factor of the keyword in the sensitive information in the initial document data.

[0115] Optionally, the target document generation module 204 is further configured to obtain access type and index information of the key data corresponding to the initial document data according to the key data corresponding to the initial document data; Obtain target operation items in the operation data of sensitive information; Target document data is generated based on the initial document data, target operation items in the sensitive information operation data, key data corresponding to the initial document data, and access type and index information of the key data corresponding to the initial document data.

[0116] Optionally, the target document generation module 204 is further configured to obtain the position of the key data in the initial document data through the index position in the index information of the key data corresponding to the initial document data; By defining a dictionary to configure the target operation item in the operation data of sensitive information and the sensitive information processing rules corresponding to the target operation item; Detect the key data at the location of the key data in the initial document data according to the access type of the key data corresponding to the initial document data, and obtain sensitive information in the key data; Obtaining a target operation item in the operation data of sensitive information in the key data, and processing the sensitive information using a sensitive information processing rule corresponding to the target operation item to obtain processed key data; The key data corresponding to the initial document data is updated with the processed key data to generate target document data.

[0117] Optionally, the information extraction module 202 is further used to obtain a user ID through fingerprint information; The initial document data is traversed by user ID, and a preset function is used to search for sensitive information in the initial document data.

[0118] In the description provided herein, algorithms and displays are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the examples of the present invention. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the present invention is not directed to any specific programming language either. It should be understood that various programming languages ​​can be utilized to implement the content of the present invention described herein, and the above description of specific languages ​​is for the purpose of disclosing the preferred embodiment of the present invention.

[0119] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0120] Those skilled in the art will appreciate that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices described in the embodiment, or alternatively may be located in one or more devices different from the devices in the examples. The modules in the foregoing examples may be combined into one module or may be divided into multiple submodules.

[0121] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition they can be divided into multiple submodules or subunits or subassemblies. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification and all processes or units of any method or device disclosed in this manner can be combined in any combination. Unless otherwise clearly stated, each feature disclosed in this specification can be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0122] In addition, some of the embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform the functions. Therefore, a processor with necessary instructions for implementing the method or method elements forms a device for implementing the method or method elements. In addition, the elements described herein of the device embodiments are examples of devices for implementing the functions performed by the elements for the purpose of implementing the invention.

[0123] As used herein, unless otherwise specified, the use of ordinal numbers "first," "second," "third," etc. to describe common objects merely indicates that different instances of similar objects are involved, and is not intended to imply that the objects so described must have a given order in time, space, order, or in any other manner.

[0124] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and teaching purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the attached description. The disclosure of the present invention is illustrative rather than restrictive with respect to the scope of the present invention, and the scope of the present invention is defined by the attached description.

Claims

1. A data leakage prevention method, characterized in that: include: Get initial document data; Acquire sensitive information in the initial document data based on the fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of a sensitive level; generating key data corresponding to the initial document data based on the sensitive information in the initial document data and the access rights of the initial document data; Target document data is generated based on the initial document data, key data corresponding to the initial document data, and operation data of the sensitive information.

2. The data leakage prevention method according to claim 1, characterized in that: The generating key data corresponding to the initial document data based on the sensitive information in the initial document data and the access rights of the initial document data includes: Obtaining keywords from the sensitive information in the initial document data; Calculating the contribution of keywords in the sensitive information in the initial document data; Acquire the access frequency of the initial document data; Key data corresponding to the initial document data is generated based on keywords in the sensitive information in the initial document data, contribution of the keywords in the sensitive information in the initial document data, access rights of the initial document data, and access frequency of the initial document data.

3. The data leakage prevention method according to claim 2, characterized in that: The calculating the contribution of the keywords in the sensitive information in the initial document data includes: Calculate the product of the frequency square term of the keyword in the sensitive information in the initial document data and the access data volume of the keyword in the sensitive information in the initial document data to obtain a first product; Calculating the product of the score of the keyword in the sensitive information in the initial document data and the length influence factor of the keyword in the sensitive information in the initial document data to obtain a second product; Calculating the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data; Performing word segmentation processing on the initial document data and counting the total number of all words in the initial document data to obtain the total number of words in the initial document data; The contribution of the keywords in the sensitive information in the initial document data is determined based on the first product, the second product, the logarithmic transformation of the access data volume of the keywords in the sensitive information in the initial document data, and the total number of words in the initial document data.

4. The data leakage prevention method according to claim 3, characterized in that: The determining the contribution of the keyword in the sensitive information in the initial document data based on the first product, the second product, the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data, and the total number of words in the initial document data includes: Adding the first product to the second product to obtain a first sum; Subtract the first sum from the logarithmic transformation of the access data volume of the keyword in the sensitive information in the initial document data to obtain a first difference value; Obtain a first preset value and a second preset value; Calculating a ratio of the total number of words in the initial document data to the first preset value to obtain a first ratio; Calculating the sum of the first ratio and the second preset value to obtain a second sum value; The ratio of the first difference value to the second sum value is calculated to obtain the contribution of the keyword in the sensitive information in the initial document data.

5. The data leakage prevention method according to claim 4, characterized in that: The generating key data corresponding to the initial document data based on the keywords in the sensitive information in the initial document data, the contribution of the keywords in the sensitive information in the initial document data, the access rights of the initial document data, and the access frequency of the initial document data includes: configuring corresponding keyword scores for the keywords in the sensitive information in the initial document data according to the contribution of the keywords in the sensitive information in the initial document data and the logarithmic transformation of the keywords in the sensitive information in the initial document data; According to the permission level, a corresponding permission score is configured for the access permission of the initial document data; Obtaining a frequency score according to the proportion of the access frequency of the initial document data to all document data; Obtaining a total word count score according to the proportion of the total word count of the initial document data in all document data; Obtaining a criticality score of a keyword in the sensitive information in the initial document data through the keyword score, the authority score, the frequency score and the total word count score; The keywords in the sensitive information in the initial document data are sorted in descending order according to the criticality scores of the keywords in the sensitive information in the initial document data, and the critical data corresponding to the initial document data are output.

6. The data leakage prevention method according to claim 3, characterized in that: The method further comprises: Calculating the length impact factor of the keyword in the sensitive information in the initial document data; The calculating the length impact factor of the keyword in the sensitive information in the initial document data includes: Acquire the length and third preset value of the keyword in the sensitive information in the initial document data; Calculating a difference between the third preset value and the length of the keyword in the sensitive information in the initial document data to obtain a second difference value; The ratio of the length of the keyword in the sensitive information in the initial document data to the second difference is calculated to obtain a length influence factor of the keyword in the sensitive information in the initial document data.

7. The data leakage prevention method according to claim 1, characterized in that: The generating target document data based on the initial document data, the key data corresponding to the initial document data and the operation data of the sensitive information includes: According to the key data corresponding to the initial document data, acquiring access type and index information of the key data corresponding to the initial document data; Obtaining a target operation item in the operation data of the sensitive information; The target document data is generated based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data.

8. The data leakage prevention method according to claim 7, characterized in that: The generating the target document data based on the initial document data, the target operation item in the operation data of the sensitive information, the key data corresponding to the initial document data, and the access type and index information of the key data corresponding to the initial document data includes: Acquiring the position of the key data in the initial document data through the index position in the index information of the key data corresponding to the initial document data; Configure the target operation item in the operation data of the sensitive information and the sensitive information processing rule corresponding to the target operation item by defining a dictionary; Detecting the key data at the location of the key data in the initial document data according to the access type of the key data corresponding to the initial document data, and acquiring sensitive information in the key data; Obtaining a target operation item in the operation data of the sensitive information in the key data, and processing the sensitive information using a sensitive information processing rule corresponding to the target operation item to obtain processed key data; The key data corresponding to the initial document data is updated using the processed key data to generate the target document data.

9. The data leakage prevention method according to claim 1, characterized in that: The obtaining of sensitive information in the initial document data based on fingerprint information includes: Obtaining a user ID through the fingerprint information; The initial document data is traversed through the user ID, and a preset function is used to search for sensitive information in the initial document data.

10. A data leakage prevention device, characterized in that: include: A data acquisition module, used to acquire initial document data; An information extraction module, used for acquiring sensitive information in the initial document data based on fingerprint information, wherein the initial document data is configured with different levels, and the sensitive information in the initial document data is data of a sensitive level; A key data generating module, configured to generate key data corresponding to the initial document data based on sensitive information in the initial document data and access rights to the initial document data; The target document generation module is used to generate target document data based on the initial document data, key data corresponding to the initial document data and operation data of the sensitive information.

Citation Information

Patent Citations

  • Data leakage preventing method and system, terminal and medium

    CN108734026A

  • Remodification, identification and alarm system and method for sensitive archives

    CN119046933A

  • Archive management system and method based on data analysis

    CN119830308A

  • Method and Systems for Virtual File Storage and Encryption

    US20180248887A1